GE-ViP: Graph-enhanced Vision Pipeline for Weakly Supervised Histopathology Slide Analysis
Maisha Rahman*, Jawad Ibn Ahad*, Md. Mehedi Hassan, Golam Moursalin, Sifat Momen
✅ Published in Neurocomputing 679 (2026) 133230 (Elsevier, Q1)
📄 Paper: ScienceDirect
DOI: 10.1016/j.neucom.2026.133230
Received: 26 November 2025 · Revised: 28 January 2026 · Accepted: 4 March 2026

Abstract
Whole Slide Images (WSIs) provide rich structural detail for cancer diagnosis, but their gigapixel scale and the lack of pixel-level annotations make automated analysis challenging. Weakly supervised learning enables slide-level training without exhaustive manual labels. However, many convolutional neural network (CNN) and multiple instance learning (MIL)-based pipelines treat patches independently, overlooking the spatial structure, while existing graph-based methods often rely on large, static graphs that are difficult to scale and sensitive to noise. We introduce GE-ViP, a lightweight Graph-Enhanced Vision Pipeline designed to overcome these limitations through a carefully integrated workflow. GE-ViP first selects diagnostically relevant tissue regions using a combination of SAM-derived tissue masks, entropy-based texture measures, and staining-related cues, effectively removing background regions and low-information patches. A Vision Transformer then encodes the selected patches to capture both fine-grained cellular features and broader tissue patterns. Next, patches are connected using a compact TriFusion-Planar graph that integrates spatial proximity, visual similarity, and semantic relationships, yielding a coherent representation of tissue organization. An edge-conditioned Graph Neural Network (GNN) processes this graph to produce slide-level predictions under weak supervision. Unlike prior pipelines that treat patch selection, feature extraction, and graph construction as independent stages, GE-ViP unifies these components into a single framework, improving efficiency, robustness, and interpretability, supported by feature-level explanations via GraphLIME.
Introduction
Histopathological analysis is a key step in cancer diagnosis, where pathologists visually inspect thin tissue slides under a microscope. Whole-slide imaging (WSI) digitizes glass slides into gigapixel images (~8×10⁴ to 2×10⁵ pixels per side), enabling efficient storage, remote access, and computational analysis. However, constructing reliable ML models requires dense patch-level annotations, which are costly and impractical at WSI scale. While this challenge is commonly addressed using MIL under slide-level supervision, existing pipelines often treat patch selection, feature extraction, and contextual aggregation as separate stages.
Key limitations of prior work:
- CNN+MIL (ABMIL, CLAM): treat patches as orderless set, losing spatial context; attention weights can be noisy
- Transformer-based MIL (TransMIL, HIPT): improve long-range reasoning but suffer from quadratic cost of self-attention on large WSIs
- Graph-based MIL (Patch-GCN, WiKG): introduce relational modeling but rely on fixed/heuristic adjacency, sensitive to tissue heterogeneity
- Post-hoc explainability (GNNExplainer, PGExplainer): computationally expensive and unstable on large heterogeneous WSI graphs
GE-ViP addresses all these limitations through a unified, lightweight pipeline with built-in interpretability.
🎯 Key Contributions
- Efficient weakly supervised pipeline — Semantic-aware patch selection using SAM masks, entropy, and saturation cues retains only 0.5k–1.5k patches per WSI (vs. 200k–300k for dense tiling), reducing preprocessing time to 10–20 min/WSI and memory to <0.2 GB.
- TriFusion-Planar graph with edge-aware GNN — A compact k-NN graph fusing: (1) planar Delaunay spatial neighbors, (2) spatial k-NN in Euclidean space, (3) appearance-based mutual k-NN, (4) semantic-based mutual k-NN. Edge-conditioned GNN with 17-dimensional edge attributes learns tissue structure from slide-level labels only.
- Multi-level interpretability — GraphLIME for node feature importance, slide-level saliency (node dropout effect), patch-level ViT rollout attention maps, subgraph visualization, and t-SNE for embedding analysis.
- Resource efficiency — 1.44M parameters, 5s per epoch (single NVIDIA RTX 4090) — the smallest and fastest among graph-based WSI methods.
Methodology
Pipeline Overview

Figure 2: A WSI is divided into patches and filtered using entropy, saturation, and SAM-based tissue masks to retain informative regions. Selected patches are encoded using a shared Vision Transformer to extract multi-scale visual features, which are combined with patch metadata to form graph nodes. GE-ViP constructs a tissue graph by fusing spatial, geometric, embedding-based, and semantic neighborhood relationships, followed by edge pruning to retain informative connections. A GNN performs message passing to produce slide-level diagnostic predictions.
Stage 1: Patch Selection
A WSI I ∈ R^{H×W×C} is divided into non-overlapping 256×256 patches at 20× magnification (~70–100k patches/slide). Three-stage filtering:
- (a) Entropy filtering: patch entropy φ^(ent) ≥ τ_ent (threshold 4.0); ensures textural richness
- (b) Saturation filtering: mean HSV saturation φ^(sat) ≥ τ_sat (threshold 0.05); eliminates poorly stained regions
(c) Semantic filtering: SAM tissue coverage score σᵢ = Mᵢ / pᵢ ≥ τ_sem (threshold 0.5)
Final selection rule: \(p_i \in \mathcal{P}^* \iff \phi_i^{(\text{ent})} \geq \tau_{\text{ent}} \wedge \phi_i^{(\text{sat})} \geq \tau_{\text{sat}} \wedge \sigma_i \geq \tau_{\text{sem}}\)
Patches are categorized into: (i) low-information benign, (ii) high-information benign, (iii) high-information suspected cancer.
Stage 2: Multi-Scale Feature Extraction
Each selected patch is processed at two spatial scales with shared ViT-Tiny/16 (224) encoder Φ:
- Standard view: Resize(224,224) → f^(1) ∈ R^192
- Context view: Resize(384,384) → CenterCrop(224,224) → f^(2) ∈ R^192
Multi-scale embedding: f_i = [f^(1)_i ‖ f^(2)_i] ∈ R^384
Semantic descriptor: m_i = [aᵢ, ρᵢ, σᵢ] ∈ R^3 (mask area, local density, semantic score)
Node feature: x_i = [f_i ‖ m_i] ∈ R^387
Stage 3: TriFusion-Planar Graph Construction
Four complementary neighborhoods are fused into slide graph G_S = (V, E):
- Planar Delaunay N^del(i) — spatial adjacency via Delaunay triangulation; bounded average degree; density-independent → geometric stability backbone
- Spatial k-NN N^xy(i) — Euclidean distance in (x,y) coordinate space
- Appearance mutual k-NN N^emb(i) — cosine similarity of visual embeddings fᵢ
- Semantic mutual k-NN N^sem(i) — similarity of tissue descriptors mᵢ
Union: U(i) = N^del ∪ N^xy ∪ N^emb ∪ N^sem, then symmetrized and degree-regularized.

Figure 3: Illustration of TriFusion-Planar graph construction. From left: (a) Delaunay triangulation over patch coordinates forms the geometric backbone; (b) spatial k-NN connects nearby patches in coordinate space; (c) appearance mutual k-NN connects visually similar patches; (d) semantic mutual k-NN connects patches of the same tissue prototype. All four neighborhoods are fused and pruned using the weighted edge scoring function.
Score-based edge pruning: \(s_{ij} = 0.45\, d_{\text{feat}}(i,j) + 0.30\, d_{\text{spat}}(i,j) + 0.25\, d_{\text{sem}}(i,j)\)
Each retained edge carries a 17-dimensional attribute vector g_ij encoding: appearance dissimilarity d^emb, spatial distance r, mask-area difference Δa, orientation (sin θ, cos θ), semantic dissimilarity d^sem, density difference Δρ, 6 RBF encodings of spatial distance, and 4 binary neighborhood source flags (1^del, 1^xy, 1^emb, 1^sem). During training, projected to compact 5-D [z(d^emb), z(r), z(Δa), sin θ, cos θ].
Stage 4: Edge-Conditioned GNN Training

Figure 4: Edge-conditioned GNN architecture. Each layer applies residual GINE message passing conditioned on 5-D edge attributes. After 4 layers, a multi-head readout (Set2Set + attention aggregation + mean pooling) produces a fixed-length slide embedding, which is classified by a temperature-scaled cosine head.
Architecture:
- 4-layer residual edge-conditioned GINE with DropPath and DropFeature
- Hidden dimension: 256
- Multi-head readout: [Set2Set ‖ AttnAgg ‖ Mean] ∈ R^{4d_hid}
- Temperature-scaled cosine classifier: ℓ = τ · (h/‖h‖) · W_c^T
Training setup:
- 120 epochs, AdamW (lr=3×10⁻⁴, weight decay=6×10⁻²), cosine warmup, mixed precision, gradient clipping
- Robustness: node noise σ=0.006, DropEdge 0.18, DropNode 0.14, DropFeature 0.18
- Loss: LDAM-DRW (class imbalance) + R-Drop λ=0.07 (consistency) + Manifold Mixup α=0.35
- EMA weight averaging (μ=0.999)
- Evaluation: 12 test-time augmentation views (node dropout p=0.084, edge dropout p=0.108)
- Early stopping: macro-F1 on validation, patience 30 epochs
- Data splits: 70:30 patient-level stratified + 10-fold patient-level cross-validation
📊 Results
Table 2: TCGA Cohort Characteristics
| Dataset | Subtype | Stage I | Stage II | Stage III | Stage IV | #WSIs | WSI Size | Disease | Task |
|---|---|---|---|---|---|---|---|---|---|
| TCGA-ESCA | Adenocarcinoma (1) | 30 | 56 | 75 | 19 | 375 | ~80k–120k px | Esophageal carcinoma | Subtype & Stage |
| Squamous cell (2) | 13 | 115 | 55 | 12 | |||||
| TCGA-KIDNEY | Chromophobe (1) | 129 | 109 | 62 | 26 | 1233 | ~100k–150k px | Renal cell carcinoma (KIRC, KIRP, KICH) | Subtype & Stage |
| Clear cell (2) | 270 | 58 | 124 | 83 | |||||
| Papillary (3) | 231 | 44 | 70 | 27 | |||||
| TCGA-LUNG | Squamous (1) | 644 | 378 | 232 | 20 | 2121 | ~100k–200k px | Lung cancer (LUAD, LUSC) | Subtype & Stage |
| Adenocarcinoma (2) | 464 | 200 | 141 | 42 |
Table 3: Cancer Type Classification (mean ± std, 10-fold CV)
Bold = best performance; underline = second best.
| Method | TCGA-ESCA Acc | TCGA-ESCA AUC | TCGA-ESCA F1 | TCGA-KIDNEY Acc | TCGA-KIDNEY AUC | TCGA-KIDNEY F1 | TCGA-LUNG Acc | TCGA-LUNG AUC | TCGA-LUNG F1 |
|---|---|---|---|---|---|---|---|---|---|
| ABMIL | 85.61±2.00 | 90.74±0.90 | 85.70±1.93 | 94.97±0.85 | 98.69±0.65 | 94.98±0.88 | 80.67±2.37 | 84.80±3.77 | 80.35±2.76 |
| CLAM-SB | 88.27±2.30 | 93.56±1.40 | 88.28±2.28 | 96.68±0.93 | 99.49±0.20 | 96.67±0.94 | 82.89±1.54 | 89.25±1.87 | 82.86±1.52 |
| CLAM-MB | 87.21±1.05 | 92.83±1.64 | 87.27±2.03 | 96.92±0.42 | 99.55±0.08 | 96.91±0.42 | 80.81±1.08 | 88.66±1.67 | 80.77±0.98 |
| DSMIL | 87.20±0.82 | 91.91±1.34 | 87.25±0.79 | 96.19±1.31 | 99.28±0.26 | 96.17±1.31 | 81.80±1.20 | 86.30±1.60 | 81.78±1.27 |
| TransMIL | 86.70±2.04 | 92.52±0.66 | 86.81±1.94 | 95.38±1.47 | 99.22±0.51 | 95.36±1.49 | 78.31±2.37 | 85.89±3.48 | 78.03±2.66 |
| DTFD-MIL | 88.81±2.49 | 93.79±2.22 | 88.88±2.43 | 96.27±1.00 | 99.30±0.36 | 96.26±1.01 | 82.41±0.82 | 88.70±1.45 | 82.37±0.88 |
| HIPT | 89.24±1.25 | 93.79±1.32 | 89.74±2.27 | 96.72±0.87 | 99.48±0.14 | 96.73±0.87 | 80.39±1.13 | 86.91±0.58 | 80.51±1.14 |
| GTP | 78.03±0.85 | 84.61±5.33 | 77.76±0.69 | 83.54±4.35 | 93.71±3.33 | 82.76±0.03 | 78.07±1.31 | 83.66±1.08 | 77.98±1.25 |
| Patch-GCN | 88.30±1.32 | 93.98±0.83 | 88.37±2.06 | 96.68±1.11 | 99.53±0.18 | 96.67±1.10 | 79.72±3.67 | 87.13±2.40 | 79.88±3.59 |
| WiKG | 90.37±2.14 | 95.23±2.90 | 90.40±3.13 | 97.08±0.71 | 99.65±0.12 | 97.08±0.71 | 84.02±0.72 | 90.78±1.24 | 83.93±0.64 |
| GE-ViP (Ours) | 92.13±1.08 | 96.11±2.15 | 91.32±1.18 | 97.56±0.85 | 99.61±0.05 | 97.02±0.66 | 87.55±0.76 | 91.81±1.31 | 86.33±1.02 |
GE-ViP achieves the highest Accuracy, AUC, and F1-score across all three TCGA cohorts for cancer type classification, surpassing the best prior graph method (WiKG) by +1.76% Acc on ESCA, +0.48% on KIDNEY, +3.53% on LUNG.
Table 4: Cancer Stage Prediction (mean ± std, 10-fold CV)
| Method | TCGA-ESCA Acc | TCGA-ESCA AUC | TCGA-ESCA F1 | TCGA-KIDNEY Acc | TCGA-KIDNEY AUC | TCGA-KIDNEY F1 | TCGA-LUNG Acc | TCGA-LUNG AUC | TCGA-LUNG F1 |
|---|---|---|---|---|---|---|---|---|---|
| ABMIL | 51.21±4.10 | 64.28±4.00 | 48.51±2.75 | 51.99±2.08 | 63.42±2.24 | 47.10±2.45 | 52.57±0.72 | 57.33±1.42 | 45.39±0.98 |
| CLAM-SB | 51.47±3.24 | 65.08±1.46 | 48.70±3.84 | 51.99±2.62 | 66.07±2.97 | 47.83±2.44 | 51.02±2.19 | 57.22±2.86 | 44.73±2.95 |
| CLAM-MB | 54.15±4.65 | 65.81±2.30 | 51.56±3.76 | 51.91±3.27 | 67.67±2.50 | 48.63±1.85 | 50.21±1.66 | 57.32±2.01 | 43.99±1.52 |
| DSMIL | 51.75±4.58 | 66.03±6.05 | 49.07±5.85 | 52.32±3.02 | 66.85±2.80 | 48.16±2.29 | 52.33±1.21 | 57.93±2.74 | 44.96±2.07 |
| TransMIL | 53.61±3.44 | 65.43±4.07 | 49.86±5.96 | 50.37±1.43 | 61.96±2.64 | 43.73±2.60 | 52.15±1.06 | 55.09±0.94 | 42.51±1.49 |
| DTFD-MIL | 50.68±5.29 | 65.60±5.18 | 48.21±5.26 | 51.75±2.00 | 64.03±2.66 | 45.63±3.38 | 52.52±0.78 | 57.56±1.71 | 44.88±1.76 |
| HIPT | 48.82±6.16 | 63.46±3.89 | 35.55±12.13 | 51.81±2.96 | 61.83±1.40 | 37.40±1.53 | 52.24±1.88 | 55.26±0.46 | 35.92±4.74 |
| GTP | 46.38±4.03 | 63.55±3.88 | 45.90±3.96 | 46.05±2.14 | 57.01±1.10 | 39.73±2.14 | 50.85±3.77 | 55.22±4.10 | 43.34±2.12 |
| Patch-GCN | 52.28±3.98 | 68.03±2.70 | 50.47±5.18 | 54.26±3.26 | 68.96±2.22 | 49.89±2.26 | 50.63±0.58 | 53.82±1.43 | 42.62±2.64 |
| WiKG | 57.50±3.99 | 69.96±4.28 | 55.80±4.39 | 55.49±2.12 | 69.71±1.40 | 51.23±1.43 | 52.85±0.74 | 60.34±1.37 | 47.52±1.52 |
| GE-ViP (Ours) | 62.74±3.54 | 71.56±4.01 | 61.71±3.02 | 56.25±1.86 | 68.46±1.32 | 50.14±2.03 | 54.88±0.43 | 62.12±1.09 | 48.35±1.44 |
Staging is inherently harder (subtle visual differences between stages I–IV). GE-ViP achieves the best Accuracy and AUC on ESCA (+5.24% Acc over WiKG) and LUNG (+2.03% AUC over WiKG), with consistent F1 improvements. GE-ViP shows particularly notable gains on the challenging staging problem where most existing methods struggle.
Table 5: Cross-Domain Evaluation — Camelyon-17 & TCGA-COAD
| Method | Camelyon-17 Acc | Camelyon-17 AUC | Camelyon-17 F1 | TCGA-COAD Acc | TCGA-COAD AUC | TCGA-COAD F1 |
|---|---|---|---|---|---|---|
| ABMIL | 76.33±0.32 | 81.10±1.05 | 55.20±1.09 | 86.24±1.40 | 95.38±0.24 | 84.39±1.11 |
| CLAM-SB | 74.18±1.26 | 78.56±1.53 | 57.22±5.01 | 86.67±0.96 | 93.65±0.34 | 84.00±1.06 |
| TransMIL | 78.10±1.05 | 91.30±1.34 | 61.02±1.37 | 86.45±1.95 | 95.07±0.98 | 83.70±2.46 |
| AMD-MIL | 75.55±0.32 | 89.05±1.05 | 69.01±1.09 | 85.81±2.07 | 95.38±0.15 | 84.21±1.59 |
| WiKG | 81.62±1.23 | 94.11±1.60 | 58.89±1.07 | 85.81±3.26 | 94.19±1.20 | 83.72±2.83 |
| FR-MIL | 77.15±1.74 | 84.43±1.16 | 71.55±1.51 | 84.09±2.07 | 92.38±1.11 | 80.26±2.23 |
| GE-ViP (Ours) | 81.84±0.61 | 94.70±1.03 | 69.80±2.55 | 87.11±0.15 | 95.21±1.01 | 85.28±0.55 |
Camelyon-17 (300 WSIs, lymph node breast metastasis): GE-ViP achieves 81.84% Acc and 94.70 AUC despite domain shift (different tissue type, staining, acquisition). TCGA-COAD (465 colon adenocarcinoma WSIs, 4 histological categories): 87.11% Acc with the lowest variance (±0.15).
📊 Ablation Study
Table 6: Patch Extraction Comparison
| Method | Magnif. | Patches/WSI | Avg. Time/WSI | Memory/WSI | Noise Induction |
|---|---|---|---|---|---|
| Dense grid tiling | 20× | 200k–300k | 100–180 min | 1.5–5 GB | Low |
| Dense grid tiling | 40× | 120k–250k | 80–120 min | 1–3 GB | Low |
| Uniform random sampling | 20× | 50k–70k | 40–60 min | 0.8–2 GB | Low |
| Tissue-masked random sampling | 20× | 15k–20k | 30–50 min | 0.5–1 GB | High |
| Entropy & saturation filtering | 20× | 5k–10k | 25–30 min | <0.5 GB | High |
| Semantic-aware (GE-ViP) | 20× | 0.5k–1.5k | 10–20 min | <0.2 GB | Low |
GE-ViP reduces patch count by >100× vs. dense tiling while maintaining low noise induction through semantic awareness. 10–20 min preprocessing vs. 100–180 min for dense tiling.
Table 7: Statistical Significance — TriFusion vs k-NN (p < 0.05 in bold)
| Model Comparison | Metric | Test | p | Significant? |
|---|---|---|---|---|
| TriFusion-ResNet18 vs kNN-ResNet18 | Accuracy | Wilcoxon | 0.043 | ✓ |
| F1 | Bootstrap | 0.3436 | ✗ | |
| AUC | Bootstrap | 0.1376 | ✗ | |
| TriFusion-ResNet50 vs kNN-ResNet50 | Accuracy | Wilcoxon | 0.0071 | ✓ |
| F1 | Bootstrap | 0.0788 | ✗ | |
| AUC | Bootstrap | 0.1836 | ✗ | |
| TriFusion-ConvNeXt-Tiny vs kNN-ConvNeXt-Tiny | Accuracy | Wilcoxon | 0.0373 | ✓ |
| F1 | Bootstrap | 0.0476 | ✓ | |
| AUC | Bootstrap | 0.0804 | ✗ | |
| TriFusion-ViT-Small vs kNN-ViT-Small | Accuracy | Wilcoxon | 0.0489 | ✓ |
| F1 | Bootstrap | 0.0948 | ✗ | |
| AUC | Bootstrap | 0.098 | ✗ | |
| TriFusion-ViT-Base vs kNN-ViT-Base | Accuracy | Wilcoxon | 0.0427 | ✓ |
| F1 | Bootstrap | 0.0524 | ✗ | |
| AUC | Bootstrap | 0.0344 | ✓ |
Table 8: Feature Extractor Ablation (AUC %)
| Backbone Type | Extractor | Tumor Type ESCA | Tumor Type KIDNEY | Tumor Type LUNG | Stage ESCA | Stage KIDNEY | Stage LUNG |
|---|---|---|---|---|---|---|---|
| CNN | ResNet-50 | 90.12 | 97.65 | 85.41 | 53.12 | 62.31 | 60.21 |
| ResNet-101 | 91.45 | 98.11 | 86.07 | 56.83 | 64.58 | 60.74 | |
| ResNet-152 | 91.98 | 98.47 | 86.55 | 59.94 | 62.27 | 58.18 | |
| ViT | ViT-Small/16 | 92.03 | 98.72 | 86.89 | 70.15 | 66.54 | 61.75 |
| ViT-Base/16 | 92.41 | 99.02 | 87.10 | 67.86 | 67.03 | 57.14 | |
| ViT-Tiny/16 (224) | 97.56 | 99.61 | 87.55 | 71.56 | 68.46 | 62.12 | |
| Pretrained | UNI | 93.82 | 98.34 | 86.91 | 64.27 | 61.11 | 56.84 |
ViT-Tiny/16 (224) delivers the strongest overall results across all tasks. ViT-based features consistently outperform CNN variants; ViT-Tiny outperforms even the pathology foundation model UNI, particularly for stage prediction — highlighting the benefit of task-adaptive feature representations integrated with graph-based reasoning.
Table 9: Edge Construction Strategy Ablation (AUC %)
| Dataset | Task | k-NN (cosine) | k-NN (distance) | GE-ViP (TriFusion) |
|---|---|---|---|---|
| TCGA-ESCA | Type | 94.98±2.52 | 92.36±1.11 | 96.11±2.15 |
| Stage | 64.36±4.38 | 63.56±4.69 | 71.56±4.01 | |
| TCGA-KIDNEY | Type | 99.44±0.28 | 99.47±0.24 | 99.61±0.11 |
| Stage | 68.62±1.97 | 68.01±1.77 | 68.46±1.32 | |
| TCGA-LUNG | Type | 90.62±1.40 | 88.31±2.22 | 91.81±1.31 |
| Stage | 58.72±1.58 | 57.92±1.99 | 62.12±1.09 |
TriFusion-Planar consistently outperforms both single-criterion k-NN baselines across all datasets and tasks. The most significant gains are on staging (+7.20% AUC on ESCA staging vs cosine k-NN).
Memory Efficiency Comparison
| Model | Epoch Time | Parameters |
|---|---|---|
| GE-ViP (Ours) | 5 s | 1.44M |
| WiKG | 8 s | 1.66M |
| Patch-GCN | 14 s | 20.14M |
| GTP | 29 s | 6.82M |
| ABMIL | ~3 s | — (non-graph) |
GE-ViP achieves the fastest training time (5s/epoch) and smallest parameter count (1.44M) among all graph-based WSI models — 14× fewer parameters than Patch-GCN.
Model Interpretation
GE-ViP provides multi-level interpretability:

Figure 5: Patch-level interpretability via ViT attention rollout. Top-K (K=12) patches selected by semantic score are visualized with attention heatmaps. Highlighted regions correspond to clinically meaningful structures: tumor-stroma boundaries, atypical nuclei clusters, glandular formations, stromal textures, and necrotic rims.

Figure 6: Slide-level saliency maps and subgraph visualization. Left: WSI thumbnail with patches shown as colored points scaled by node importance (confidence drop when zeroed out) — salient nodes cluster over tumor nests, invasive fronts, and gland chains. Right: Top-ranked edges drawn as blue connections, with thick bundles tracing coherent architectural motifs (tumor boundaries, dense stromal regions).

Figure 7: t-SNE visualization of learned slide embeddings. Classes form compact, well-separated clusters. Slides at cluster centers exhibit clear, canonical cancer morphology; boundary slides display ambiguous tissue patterns (useful for clinical triage). GraphLIME feature importance (inset) shows boundary/structure-sensitive ViT channels and texture/nuclear descriptors rank highest; color-only features contribute less.
- Patch-level (ViT rollout): Attention rollout on top-K (K=12) patches by semantic score → highlights tumor-stroma boundaries, atypical nuclei clusters, stromal textures, necrotic rims — clinically meaningful regions
- Slide-level saliency: Each patch shown as a point on WSI thumbnail with color/size encoding node importance (drop in argmax confidence when zeroed) → salient nodes cluster over tumor nests, invasive fronts, gland chains, stromal reaction
- Subgraph view: Top-ranked edges (by endpoint node saliency sum) displayed as blue connections → thick edge bundles trace coherent architectural motifs
- GraphLIME feature importance: Boundary/structure-sensitive ViT channels and texture/nuclear descriptors rank highest; color-only features contribute less (likely due to stain normalization independence)
- t-SNE embeddings: Slides deep inside same-class clusters show clear representative cancer morphology; boundary slides display ambiguous tissue patterns → clinical triage utility
📚 Citation
@article{rahman2026gevip,
title={GE-ViP: Graph-enhanced vision pipeline for weakly supervised histopathology slide analysis},
author={Rahman, Maisha and Ahad, Jawad Ibn and Hassan, Md. Mehedi and Moursalin, Golam and Momen, Sifat},
journal={Neurocomputing},
volume={679},
pages={133230},
year={2026},
publisher={Elsevier},
doi={10.1016/j.neucom.2026.133230}
}
