Joint Multi-Task Deep Learning with Cross-Task Attention for Simultaneous Lesion Segmentation, Detection, and Severity Grading of Diabetic Retinopathy on Fundus Images: An External Validation Study.
L’essentiel
Diabetic retinopathy (DR) severity is graded from the type and distribution of retinal lesions, yet most automated methods address grading or lesion analysis in isolation. This study develops one framework that performs lesion segmentation, lesion detection, and DR grading together, evaluated within and across datasets. A hierarchical multi-task network with a shared ResNet-50 encoder and a feature pyramid feeds three heads, in which a cross-task attention module conditions the grading representation on the predicted lesion features. The three losses are combined by homoscedastic uncertainty weighting under a progressive three-stage schedule that fits the partial-label structure of the data. The model was developed on DDR (13,673 images, 757 with lesion annotation) using stratified five-fold cross-validation over three seeds, and evaluated without tuning on the independent IDRiD dataset. Results are reported for the training, validation, and test splits, with 95 percent bootstrap confidence intervals, paired significance tests, calibration, and an ablation. On the DDR test split, the model reached a mean lesion-segmentation Dice of 0.535, a detection means average precision at intersection over union (IoU) 0.5 of 0.289, a grading accuracy of 0.823, and a quadratic weighted kappa of 0.858, with a referable-DR sensitivity of 0.928 and a vision-threatening-DR sensitivity of 0.871. On the external IDRiD set, segmentation Dice was 0.489 and grading accuracy and kappa were 0.781 and 0.821, quantifying the reduction under domain shift; the external segmentation estimates in particular rest on only 81 images and are indicative rather than definitive. In the ablation, removing the cross-task attention produced the largest single-component reduction in grading performance, while collapsing the framework to single-task training lowered it further. Jointly learning lesion localization and grading, with external validation and calibration, produces a transparent framework whose grading is accompanied by lesion-level evidence. However, external performance declines moderately (Dice: 0.535 → 0.489), and the small external segmentation sample (n = 81) limits generalizability claims. Prospective validation on larger multi-ethnic cohorts is needed before clinical use.
Synthèse détaillée
Résumé original
Diabetic retinopathy (DR) severity is graded from the type and distribution of retinal lesions, yet most automated methods address grading or lesion analysis in isolation. This study develops one framework that performs lesion segmentation, lesion detection, and DR grading together, evaluated within and across datasets. A hierarchical multi-task network with a shared ResNet-50 encoder and a feature pyramid feeds three heads, in which a cross-task attention module conditions the grading representation on the predicted lesion features. The three losses are combined by homoscedastic uncertainty weighting under a progressive three-stage schedule that fits the partial-label structure of the data. The model was developed on DDR (13,673 images, 757 with lesion annotation) using stratified five-fold cross-validation over three seeds, and evaluated without tuning on the independent IDRiD dataset. Results are reported for the training, validation, and test splits, with 95 percent bootstrap confidence intervals, paired significance tests, calibration, and an ablation. On the DDR test split, the model reached a mean lesion-segmentation Dice of 0.535, a detection means average precision at intersection over union (IoU) 0.5 of 0.289, a grading accuracy of 0.823, and a quadratic weighted kappa of 0.858, with a referable-DR sensitivity of 0.928 and a vision-threatening-DR sensitivity of 0.871. On the external IDRiD set, segmentation Dice was 0.489 and grading accuracy and kappa were 0.781 and 0.821, quantifying the reduction under domain shift; the external segmentation estimates in particular rest on only 81 images and are indicative rather than definitive. In the ablation, removing the cross-task attention produced the largest single-component reduction in grading performance, while collapsing the framework to single-task training lowered it further. Jointly learning lesion localization and grading, with external validation and calibration, produces a transparent framework whose grading is accompanied by lesion-level evidence. However, external performance declines moderately (Dice: 0.535 → 0.489), and the small external segmentation sample (n = 81) limits generalizability claims. Prospective validation on larger multi-ethnic cohorts is needed before clinical use.