This cross-sectional, population-based study (EPIC-Norfolk Eye Study, n=6,304; mean age 68 years, 57% women) compared vertical cup-disc ratio (VCDR) estimated by trained human graders (H-VCDR) versus a machine learning model (ML-VCDR) from color fundus images for detecting specialist-confirmed glaucoma.
ML-VCDR substantially outperformed human graders: right eye AUROC 90% (95% CI 88.7–91.0) vs. 81% (95% CI 78.6–82.5) for H-VCDR, and explained 35% vs. 20% of glaucoma status variance. ML-VCDR also outperformed both AutoMorph and Heidelberg Retinal Tomography (HRT).
- VCDR is a single structural measure; the ML model did not use other glaucoma-relevant features (e.g., RNFL patterns, visual field data). - The ML model was trained externally on UK Biobank pseudo-labels, raising questions about generalizability to non-European or younger populations. - Cross-sectional design cannot establish causation or assess screening program performance over time.
ML-based VCDR grading from standard fundus photos may meaningfully improve glaucoma population screening compared to human graders — clinicians and health systems designing screening programs should consider ML-assisted fundus image analysis. However, real-world implementation should account for model generalizability beyond the older, predominantly female, European cohort studied here.
Explore related topics