Deep learning shows promise for diabetic retinopathy grading
Deep learning models show strong potential for diabetic retinopathy screening and initial severity grading, particularly for identifying eyes without disease and those with vision-threatening retinopathy, although distinguishing between adjacent disease stages remains challenging, according to a systematic review and meta-analysis published in Frontiers in Endocrinology.
A team of researchers analyzed 41 diagnostic accuracy studies encompassing more than 500,000 retinal fundus images from 16 countries to evaluate the performance of deep learning algorithms using the International Clinical Diabetic Retinopathy (ICDR) severity scale. The review included a wide range of convolutional neural networks, transformer-based models, and hybrid architectures trained on both public and private datasets, according to Xin Yan from The Affiliated Hospital (Clinical College) of Xiangnan University, Chenzhou, China,
In the five-stage ICDR classification, pooled sensitivity varied considerably across disease severity. Deep learning models achieved a sensitivity of 95.2% for identifying eyes without diabetic retinopathy, 72.1% for mild nonproliferative diabetic retinopathy (NPDR), 84.3% for moderate NPDR, 75.8% for severe NPDR, and 78.8% for proliferative diabetic retinopathy (PDR). The study authors found that errors most commonly occurred between adjacent disease stages, reflecting the difficulty of distinguishing subtle differences in retinal pathology.
Performance improved when investigators evaluated simplified four-stage grading systems that merged some nonproliferative categories. Sensitivity increased to 96.9% for stage 0, 92.9% for stage 1, 92.8% for stage 2, and 88.2% for stage 3, which represented vision-threatening diabetic retinopathy in the simplified system. However, the researchers cautioned that although simplified classification improves overall diagnostic performance, it may reduce the ability to detect subtle disease progression that can influence clinical management.
The analysis suggests that deep learning systems may be particularly valuable for large-scale screening and referral triage, especially in settings with limited ophthalmology resources. High sensitivity for identifying eyes without disease and those with potentially referable disease could help streamline screening programs while reducing specialist workload. Despite these findings, the authors noted that “current technology has not yet fully matched the proficiency of clinical experts in nuanced grading, especially in differentiating between early NPDR subtypes.”
The study goes on to identify several barriers to broader clinical implementation of deep learning models for diabetic retinopathy grading. Notably, only five of the included studies performed external or multicenter validation, raising concerns about how well existing algorithms generalize across different populations, imaging devices, and clinical settings. Furthermore, nearly one in five studies carried a high risk of bias related to the reference standard because of inadequate reporting of expert grading procedures or reliance on publicly assigned image labels without independent verification.
The investigators also noted that current AI systems rely primarily on fundus photographs and vascular features, which limits their ability to detect other non-vascular pathophysiologic changes, such as retinal neurodegeneration and inflammation, Future models, the authors noted, may benefit from integrating multimodal imaging—including optical coherence tomography angiography and ultra-widefield imaging—to move beyond the limitations of the current grading framework. Improving model explainability may also increase clinician confidence in AI-assisted decisions.
Beyond technical performance, the review also highlighted several practical challenges to real-world deployment, including regulatory approval, workflow integration, medicolegal responsibility, and the lack of cost-effectiveness data. The authors noted that future research should move beyond algorithm development to prospective implementation studies evaluating clinical outcomes, health economics, and standardized reporting.
The authors concluded that standardized datasets, rigorous external validation, multimodal imaging, and unified reporting standards will be essential to support the broader clinical adoption of AI-assisted diabetic retinopathy grading and improve the precision, efficiency, and accessibility of care.
The study was supported by the Hunan Provincial Natural Science Foundation, the Hunan Provincial Education Department, Xiangnan University, and the Chenzhou Lacrimal Minimally Invasive Diagnosis and Treatment Technology Research and Development Center. The authors reported no commercial or financial relationships that could be construed as potential conflicts of interest.
AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.