Graph-enhanced deep learning for diabetic retinopathy diagnosis: A quality-aware and uncertainty-driven approach
Zarin Akter, Jawad Ibn Ahad, Md. Mutasim Farhan, Riasat Khan
β Published in PLOS Computational Biology (Q1, Open Access)
π Paper: PLOS Computational Biology
DOI: 10.1371/journal.pcbi.1013745
π» Code: github.com/mfar201/diabetic_retinopathy_classification_gcn
Received: April 02, 2025 Β· Accepted: November 13, 2025 Β· Published: December 5, 2025
Abstract
Diabetic retinopathy (DR) is a leading cause of vision impairment, significantly impacting working-class populations and necessitating accurate early diagnosis. Traditional DR classification relies on CNN-based models and extensive preprocessing. We propose a novel approach leveraging pre-trained models for feature extraction, followed by Graph Convolutional Networks (GCNs) for refined embedding representation. The extracted feature vectors are structured as a graph, where GCN enhances embeddings before classification. The model incorporates quality assessment (QA) by predicting a confidence score through a dedicated fully connected layer trained with binary cross-entropy loss, and uncertainty estimation (UE) by calculating variance across multiple stochastic passes. Evaluated on APTOS2019, Messidor-2, and EyePACS, the framework achieves superior performance over state-of-the-art methods: 98.45% accuracy (MobileViT, APTOS2019), 94.90% accuracy (DenseNet-169, Messidor-2), and 97.38% accuracy (DenseNet-169, EyePACS) β all without requiring intensive image preprocessing.
Introduction
Figure 1: Traditional DR classification (top) relies on CNN-based models with extensive preprocessing on raw fundus images. The proposed strategy (bottom) utilizes pre-trained models for feature extraction (FE). The generated feature vectors (FV) are refined using a GCN and subsequently leveraged for classification, quality assessment (QA), and uncertainty estimation (UE).
Diabetic retinopathy progresses in four stages: (a) Mild NPDR β microaneurysms in retinal blood vessels; (b) Moderate NPDR β increased microaneurysms, hemorrhages, hard exudates; (c) Severe NPDR β body signals abnormal vessel formation; (d) Proliferative DR (PDR) β neovascularization causing fragile leaking vessels and potential blindness. 537 million adults worldwide live with diabetes (IDF); 22.27% suffer DR (2021).
Existing deep learning models excel at binary DR classification but struggle with reliable multiclass grading. They also depend on heavy preprocessing (CLAHE, Ben Graham normalization) and lack quality awareness and uncertainty quantification. We address all these gaps in a unified GCN-based framework.
π― Key Contributions
- Preprocessing-free pipeline β GCN framework achieves SOTA performance directly on fundus images with only basic resizing and normalization, without CLAHE or Ben Graham preprocessing.
- Graph-Structured Feature Refinement β Fundus image features structured as a graph G = (V, E); GCN aggregates spatial and semantic neighborhood information to enhance embeddings.
- Quality Assessment (QA) β Fully connected layer predicts confidence score qΜ = Ο(W_QA h^L + b_QA) β [0,1], trained with binary cross-entropy loss.
- Uncertainty Estimation (UE) β T=10 Monte Carlo dropout passes compute prediction variance ΟΒ² = (1/T)Ξ£(Ε·^(t) β Θ³)Β²; high variance flags uncertain cases for clinical review.
- Grad-CAM Interpretability β Heatmaps demonstrate clinically relevant focus on retinal pathologies (microaneurysms, hemorrhages, neovascularization).
- 15 backbone architectures evaluated β CNN (DenseNet, ResNet, Inception, EfficientNet, Xception) and Transformer (ViT, Swin, DeiT, MobileViT).
Methodology
Figure 2: Model architecture. Dataset D undergoes basic preprocessing (resize to 224Γ224, transform, rotation). FE function f processes each sample x β D to generate FV z β R^d, refined to R^d via Global Average Pooling (GAP). A graph G = (V, E) is constructed with nodes corresponding to FVs; edge distance is computed from spatial distance d_sp and semantic distance d_se. Two GCN layers refine embeddings h^(1) β R^512 β h^(2) β R^256. Two FCLs produce: Ε· β R^{256Γ5} (classification) and qΜ β [0,1] (quality assessment). Total loss: L_total = L_cls + Ξ»L_q.
Problem Formulation
Input dataset D of retinal images x_i β R^{HΓW}. A backbone f extracts feature vector z = f(x) β R^d. Graph G = (V, E) is constructed where nodes = features z, edges = spatial + semantic distances. GCN refines embeddings h. Softmax classifier: Ε· = softmax(W_cls h + b_cls) β R^5.
Graph Construction
Combined edge distance between nodes i and j: \(d_{\text{comb}}(i,j) = \beta \cdot d_{\text{sp}}(i,j) + (1-\beta) \cdot d_{\text{se}}(i,j), \quad \beta \in [0,1]\)
Parameters: k=4 nearest neighbors, radius=0.1, Ξ²=0.5.
GCN Layer Update Rule
\[\mathbf{h}_i^{(l+1)} = \sigma\!\left(\sum_{j \in \mathcal{N}(i)} \frac{1}{\sqrt{\deg(i)\deg(j)}} \mathbf{W}_l \mathbf{h}_j^{(l)} + \mathbf{b}_l\right)\]Two GCN layers: R^1024 β R^512 β R^256.
Quality Assessment & Uncertainty Estimation
QA loss: \(\mathcal{L}_q = \text{BCE}(\hat{q}, q) = -q\log(\hat{q}) - (1-q)\log(1-\hat{q})\)
Uncertainty (T=10 MC dropout passes): \(\bar{y} = \frac{1}{T}\sum_{t=1}^T \hat{y}^{(t)}, \qquad \sigma = \sqrt{\frac{1}{T}\sum_{t=1}^T (\hat{y}^{(t)} - \bar{y})^2}\)
Dropout rates: p=0.3 (classifier head), p=0.2 (GCN layers). Total loss: L_total = L_cls + Ξ»L_q (Ξ»=0.1).
π Datasets
Table 2: Dataset Statistics (5 DR Severity Classes)
| Class | DR Grade | APTOS2019 | Messidor-2 | EyePACS |
|---|---|---|---|---|
| Class-0 | No DR | 1,805 | 1,017 | 25,810 |
| Class-1 | Mild NPDR | 999 | 270 | 2,443 |
| Class-2 | Moderate NPDR | 370 | 347 | 5,292 |
| Class-3 | Severe NPDR | 295 | 75 | 873 |
| Class-4 | PDR | 193 | 35 | 708 |
| Total | Β | 3,662 | 1,748 | 35,126 |
Dataset split: 70% train / 15% validation / 15% test (stratified). Class imbalance handled via oversampling with Albumentations augmentation (rotations, flips, blur, brightness/contrast).
π Experimental Setup
Table 3: Hyperparameter Values
| Category | Hyperparameter | Value |
|---|---|---|
| Training | Epochs | 50 |
| Β | Batch Size | 32 |
| Β | Learning Rate | 5e-5 |
| Β | Optimizer | AdamW |
| Β | Weight Decay | 0.01 |
| Β | LR Scheduler | ReduceLROnPlateau |
| Β | Scheduler Patience | 7 |
| Β | Early Stopping Patience | 15 |
| Graph Construction | Neighbors (k) | 4 |
| Β | Radius | 0.1 |
| Β | Feature Weight (Ξ²) | 0.5 |
| Uncertainty | MC Dropout Samples (T) | 10 |
| Β | Inference Dropout Rate | 0.3 |
| GCN | Input Dimension | 1024 |
| Β | Hidden Dimensions | [512, 256] |
| Β | Output Dimension | 1024 |
| Β | Dropout Rate | 0.2 |
| Loss | Classification Weight | 1.0 |
| Β | QA Loss Weight (Ξ») | 0.1 |
Table 4: Model-Specific Training Parameters (15 Backbones)
| Architecture | Backbone | Model Size (MB) | Parameters | SIIT (ms) | Time/Epoch (s) |
|---|---|---|---|---|---|
| CNN | ResNet50 | 295.27 | 25,755,462 | 16.59 | 88.95 |
| Β | ResNet101 | 513.07 | 44,747,590 | 19.23 | 95.61 |
| Β | ResNet152 | 692.52 | 60,391,238 | 22.82 | 1138.60 |
| Β | Xception | 458.76 | 40,017,990 | 15.26 | 1458.57 |
| Β | InceptionV3 | 275.69 | 24,032,998 | 18.88 | 93.30 |
| Β | InceptionResNetV2 | 642.58 | 56,025,510 | 36.20 | 100.40 |
| Β | EfficientNetB3 | 142.99 | 12,415,278 | 17.82 | 99.00 |
| Β | DenseNet121 | 94.20 | 8,144,518 | 20.92 | 114.47 |
| Β | DenseNet161 | 332.29 | 28,884,550 | 26.13 | 103.75 |
| Β | DenseNet169 | 165.58 | 14,335,622 | 23.34 | 97.11 |
| Β | DenseNet201 | 233.23 | 20,208,262 | 28.93 | 105.91 |
| Transformer | ViT-Base | 992.76 | 86,725,126 | 15.65 | 104.87 |
| Β | Swin-Base | 1006.84 | 87,933,886 | 24.36 | 101.43 |
| Β | DeiT-Base | 992.77 | 86,726,662 | 13.49 | 91.84 |
| Β | MobileViT | 28.60 | 2,463,030 | 19.22 | 98.91 |
SIIT = Single Image Inference Time. Hardware: NVIDIA GeForce RTX 4070 (12GB VRAM), PyTorch.
π Results
Figure 3: Normalized Confusion Matrices
Figure 3: Normalized confusion matrices for multiclass DR classification on APTOS2019 and Messidor-2 under three preprocessing conditions: No Preprocessing (left), CLAHE (middle), Ben-Graham (right). MobileViT on APTOS2019 shows excellent performance with minimal misclassification (AP=1.00). DenseNet-169 achieves high accuracy on Messidor-2.
Figure 4: Precision-Recall Curves
Figure 4: Precision-recall curves for all 5 DR classes on APTOS2019 and Messidor-2 under three preprocessing conditions. Our pipeline (No Preprocessing) achieves AP=1.00 across all classes on APTOS2019 with MobileViT.
Table 5: Full SOTA Comparison (APTOS2019 & Messidor-2)
Legend: β CNN Β· β΄ Transformer Β· β£ CLAHE preprocessing Β· β¦ Ben-Graham preprocessing Β· β‘ Our pipeline (no preprocessing)
Prior State-of-the-Art Methods
| Arch | Ref | Backbone | APTOS Acc | APTOS F1 | APTOS AUROC | APTOS Kappa | Messidor Acc | Messidor F1 | Messidor AUROC | Messidor Kappa |
|---|---|---|---|---|---|---|---|---|---|---|
| β | [6] | ResNet50 | 84.80 | 84.30 | β | 90.90 | 67.10 | 65.50 | β | 66.30 |
| β | [6] | DenseNet121 | 85.50 | 84.90 | β | 90.60 | 68.00 | 66.00 | β | 67.30 |
| β | [17] | DenseNet121 | 97.30 | β | β | β | β | β | β | β |
| β | [18] | RSG-Net | β | β | β | β | 99.36 | 99.40 | 99.98 | β |
| β | [21] | ResNet50 | 85.65 | β | 89.00 | β | β | β | β | β |
| β | [38] | DenseNet201 | 91.62 | 91.52 | β | β | 85.79 | 85.08 | β | β |
| β | [38] | MobileNetV2 | 93.09 | 93.53 | β | β | 83.81 | 85.23 | β | β |
| β | [40] | DenseNet121 | 97.68 | 97.00 | 96.20 | 98.50 | β | β | β | β |
| β | [43] | Xception | 84.36 | 70.49 | 93.82 | β | 74.21 | 55.18 | 87.26 | β |
| β΄ | [44] | Swin-Base | 88.70 | 88.70 | β | β | 83.12 | 83.12 | β | β |
Our Pipeline β‘ β No Preprocessing (Primary Results)
| Arch | Backbone | APTOS Acc | APTOS F1 | APTOS AUROC | APTOS AUPR | APTOS Kappa | Messidor Acc | Messidor F1 | Messidor AUROC | Messidor AUPR | Messidor Kappa |
|---|---|---|---|---|---|---|---|---|---|---|---|
| β | ResNet50 | 93.95 | 94.02 | 99.70 | 98.98 | 92.44 | 70.46 | 67.28 | 92.55 | 78.90 | 63.07 |
| β | ResNet101 | 84.65 | 84.04 | 97.79 | 93.54 | 80.81 | 40.00 | 22.86 | 79.60 | 54.07 | 25.00 |
| β | ResNet152 | 92.18 | 92.13 | 99.25 | 97.67 | 90.22 | 79.87 | 80.07 | 96.53 | 90.61 | 74.84 |
| β | Xception | 95.28 | 95.26 | 99.59 | 98.67 | 94.10 | 93.07 | 93.05 | 99.14 | 97.34 | 91.34 |
| β | InceptionV3 | 96.97 | 96.97 | 99.82 | 99.38 | 96.22 | 94.12 | 94.08 | 99.49 | 98.35 | 92.65 |
| β | InceptionResNetV2 | 94.98 | 94.95 | 99.41 | 97.97 | 93.73 | 94.38 | 94.37 | 99.37 | 97.95 | 92.97 |
| β | EfficientNetB3 | 96.83 | 96.82 | 99.83 | 99.40 | 96.03 | 94.12 | 94.12 | 99.52 | 98.42 | 92.65 |
| β | DenseNet121 | 93.87 | 93.80 | 99.44 | 98.12 | 92.34 | 87.58 | 87.59 | 98.65 | 96.06 | 84.48 |
| β | DenseNet161 | 96.24 | 96.23 | 99.70 | 99.02 | 95.30 | 94.38 | 94.39 | 99.52 | 98.41 | 92.97 |
| β | DenseNet169 | 96.75 | 96.74 | 99.81 | 99.37 | 95.94 | 94.90 | 94.87 | 99.50 | 98.42 | 93.63 |
| β | DenseNet201 | 96.61 | 96.60 | 99.84 | 99.43 | 95.76 | 92.94 | 92.92 | 99.35 | 97.90 | 91.18 |
| β΄ | ViT-Base | 95.57 | 95.55 | 99.62 | 98.79 | 94.46 | 88.63 | 88.66 | 98.06 | 94.17 | 85.78 |
| β΄ | Swin-Base | 96.90 | 96.89 | 99.83 | 99.44 | 96.13 | 91.24 | 91.23 | 99.04 | 96.85 | 89.05 |
| β΄ | DeiT-Base | 95.72 | 95.71 | 99.71 | 99.07 | 94.65 | 90.20 | 90.17 | 98.91 | 96.52 | 87.75 |
| β΄ | MobileViT | 98.45 | 98.45 | 99.94 | 99.81 | 98.06 | 92.03 | 92.02 | 99.25 | 97.45 | 90.03 |
CLAHE Preprocessing β£ (Comparison)
| Arch | Backbone | APTOS Acc | APTOS Kappa | Messidor Acc | Messidor Kappa |
|---|---|---|---|---|---|
| β | MobileViT | 98.01 | 97.51 | 91.37 | 89.22 |
| β | DenseNet169 | 96.68 | 95.85 | 93.46 | 91.83 |
| β | EfficientNetB3 | 96.97 | 96.22 | 92.94 | 91.18 |
| β΄ | Swin-Base | 96.90 | 96.13 | 91.24 | 89.05 |
Ben-Graham Preprocessing β¦ (Comparison)
| Arch | Backbone | APTOS Acc | APTOS Kappa | Messidor Acc | Messidor Kappa |
|---|---|---|---|---|---|
| β | MobileViT | 93.95 | 92.44 | 92.55 | 90.69 |
| β | DenseNet169 | 94.83 | 93.54 | 95.03 | 93.79 |
| β | EfficientNetB3 | 93.21 | 91.51 | 91.50 | 89.38 |
| β΄ | Swin-Base | 95.42 | 94.28 | 90.59 | 88.24 |
Our pipeline (β‘) with NO preprocessing achieves the best overall results. MobileViT β‘ reaches 98.45% Acc and 98.06% Kappa on APTOS2019 β surpassing CLAHE (98.01%) and Ben-Graham (93.95%) variants. DenseNet-169 β‘ achieves 94.90% Acc on Messidor-2. Extensive preprocessing is unnecessary and often harmful.
Table 6: External Validation β EyePACS Dataset
| Initially Trained On | Backbone | EyePACS Acc | EyePACS F1 | EyePACS AUROC | EyePACS AUPR | EyePACS Kappa |
|---|---|---|---|---|---|---|
| Messidor-2 | DenseNet-169 | 97.38% | 97.37% | 99.83% | 99.39% | 96.72% |
| APTOS2019 | MobileViT | 96.02% | 96.02% | 99.60% | 98.69% | 95.03% |
Strong cross-dataset generalization: DenseNet-169 pretrained on Messidor-2, fine-tuned on EyePACS (35,126 images), achieves 97.38% accuracy β demonstrating the frameworkβs robustness across diverse imaging conditions, demographics, and grading variations.
Explainability β Grad-CAM
Figure 5: Grad-CAM heatmaps for retinal fundus images across 5 DR severity levels. Each row: original image, Grad-CAM heatmap, prediction. Red regions indicate high classification influence. (a) No preprocessing (our approach). (b) CLAHE preprocessing. Stage-by-stage clinical interpretation:
- No DR (Class 0): Diffuse, unfocused activations β confirming absence of pathological markers
- Mild NPDR (Class 1): Small punctate activation areas β highlighting microaneurysms (earliest DR signs)
- Moderate NPDR (Class 2): Larger, more pronounced regions β aligned with dot/blot hemorrhages and hard exudates
- Severe NPDR & PDR (Class 3β4): Large intense activation areas β focused on retinal hemorrhages and neovascularization
π Ablation Study
Figure 6 (a): Impact of class balancing strategies on APTOS2019 (MobileViT). Compared: original imbalanced dataset (OgD), Compute Class Weight (OgD WC), Weighted Random Sampler (OgD RS). Our balanced oversampling approach (black bar) consistently achieves highest Accuracy, F1, AUROC, AUPR, and Kappa. (b) AdamW vs SGD optimizer: AdamW outperforms SGD across all metrics, demonstrating better generalization and robustness.
π Citation
@article{akter2025graph,
title={Graph-enhanced deep learning for diabetic retinopathy diagnosis: A quality-aware and uncertainty-driven approach},
author={Akter, Zarin and Ahad, Jawad Ibn and Farhan, Md. Mutasim and Khan, Riasat},
journal={PLOS Computational Biology},
volume={21},
number={12},
pages={e1013745},
year={2025},
publisher={Public Library of Science},
doi={10.1371/journal.pcbi.1013745}
}
