WHAT THE STUDY ACTUALLY SAYS

AI matched or beat eye doctors at spotting macular degeneration — on retrospective data

Two 2026 meta-analyses pooled deep-learning systems reading retinal images for AMD. The accuracy looks near-perfect, but almost all of it comes from algorithms tested on their own data.

Pooled sensitivity for detecting AMD: deep learning vs senior ophthalmologistsDeep-learning algorithms: 0.98; Senior ophthalmologists: 0.7500.51Deep-learning algorithms0.98Senior ophthalmologists0.75
Pooled sensitivity for detecting AMD: deep learning vs senior ophthalmologists
GroupValue (value)
Deep-learning algorithms0.98
Senior ophthalmologists0.75
Pooled sensitivity for detecting AMD: deep learning vs senior ophthalmologists Bivariate random-effects meta-analysis of head-to-head studies; 28 studies, 77,485 samples. Source: Journal of Medical Internet Research

Deep-learning systems that read retinal images can flag age-related macular degeneration (AMD) — the leading cause of irreversible central vision loss in older adults — and two 2026 meta-analyses find they match or beat human graders on paper [s1][s2]. Both also warn that the headline numbers come almost entirely from retrospective studies testing algorithms on their own data, so they describe potential rather than proven screening performance [s1][s2].

The task the models are doing

AMD is detected from retinal images — colour photographs of the back of the eye or optical coherence tomography (OCT) cross-sections. Two distinct jobs matter: telling AMD from a normal retina, and, within AMD, separating the "wet" neovascular form that needs prompt treatment from the slower "dry" form [s1]. For screening, the practical target is referable AMD — intermediate or advanced disease that warrants a specialist's attention [s2]. Deep learning has been proposed as a way to do this at scale, where there are far more retinas to screen than ophthalmologists to read them.

What the numbers say

The larger meta-analysis pooled 28 studies, covering 77,485 image samples for AMD detection and 28,705 for the wet-versus-dry distinction [s1]. For detecting AMD, deep learning achieved a pooled sensitivity of 0.98 (95% CI 0.96–0.99), specificity of 0.98 (95% CI 0.95–0.99), accuracy of 0.97, and an area under the curve of 1.00 [s1]. For wet versus dry AMD it reported sensitivity and specificity of 0.95 each, with an AUC of 0.99 [s1]. In head-to-head comparisons, the algorithms showed higher sensitivity than senior ophthalmologists — 0.98 against 0.75 (P < .001) — and OCT-based models performed more consistently than colour fundus photography or multimodal models [s1].

The second meta-analysis, focused on referable AMD from fundus photographs, is a useful reality check because its numbers are lower for a narrower, arguably more clinically relevant question. Across 14 studies it found a pooled sensitivity of 0.91 (95% CI 0.86–0.94) and specificity of 0.93 (95% CI 0.86–0.96), with substantial heterogeneity, and positive and negative likelihood ratios of 12.22 and 0.10 — figures that indicate strong diagnostic utility [s2]. Here the deep-learning systems showed slightly lower sensitivity but higher specificity than human graders [s2]. The two syntheses disagree in the informative way: change the endpoint from "any AMD" to "referable AMD," and the comparison with clinicians flips direction. A single accuracy figure means little without the question it answered.

Why "near-perfect" needs reading twice

The first meta-analysis is candid about why its own numbers should not be taken to the clinic. It cites high heterogeneity, wide prediction intervals, predominantly retrospective study designs and possible performance inflation from internal validation, and concludes the relative-performance findings are preliminary rather than deployment-ready [s1]. Its recommendation is explicit: deep learning should be viewed as a triage adjunct requiring local calibration, not an autonomous diagnostic replacement, and what is needed is prospective, multicentre, patient-level external validation with prespecified human comparison arms [s1]. An AUC of 1.00 is a signal to be suspicious of the test set, not to declare the problem solved.

The kind of study that is missing

One 2026 project is building toward the evidence the meta-analyses say is absent. I-SCREEN, a pan-European effort across six countries, is developing an OCT-based infrastructure to detect AMD and predict its progression through a shared-care network of community optometry practices and ophthalmology clinics — 28 community practices and 7 clinics in its screening network [s3]. Its PYRENEES arm is a prospective, cross-sectional study testing whether subclinical AMD can be picked up in optometry practices under ophthalmologist supervision via telemedicine, with suspected cases referred onward [s3]. It is a methodological framework rather than a results paper, but its prospective, community-based, referral-linked design is exactly what a retrospective accuracy meta-analysis cannot supply.

What it means, and what to watch

The 2026 evidence says deep learning reads AMD images accurately in the lab and can rival or exceed clinicians on curated data — a genuine capability, and a familiar trap. The gap is between diagnostic accuracy and demonstrated screening value: prospective performance in the messy, referral-connected settings where screening actually happens, on populations and cameras the model has not seen. Watch for whether tools like these clear that bar, or whether they join the roster of AI imaging systems adopted on internal-validation numbers alone.

For comparison, see our coverage of autonomous AI for diabetic retinopathy that did reach real-world validation, smartphone-based AI eye screening, an ophthalmology foundation model in a randomised trial, and the wider gap between AI device accuracy and patient outcomes.

Sources

Sources

  1. Performance of Deep Learning in Classifying Age-Related Macular Degeneration From Images: Systematic Review and Meta-Analysis — Journal of Medical Internet Research , June 15, 2026
  2. Performance and Clinical Utility of Deep Learning for Detecting Referable Age-Related Macular Degeneration on Fundus Photographs: A Systematic Review and Meta-Analysis — Diagnostics , February 23, 2026
  3. I-SCREEN: Development of an AI-based infrastructure for community-wide screening and prediction of progression in age-related macular degeneration providing accessible shared care — Eye , May 13, 2026
Related coverage