An AI eye copilot raised correct diagnoses from 75% to 92% in a randomised trial
EyeFM was tested as an assistant to 16 ophthalmologists screening 668 high-risk patients in China. Patients in the AI arm also followed referral advice more often.
| Group | Value (%) |
|---|---|
| Ophthalmologist with EyeFM copilot | 92.2 |
| Standard care | 75.4 |
Most medical AI evidence consists of retrospective accuracy comparisons: the model versus the clinician, on stored images, with the answer already known. A paper published in Nature Medicine on 28 August reports something rarer — a double-masked randomised controlled trial in which the model was used by clinicians during real consultations, with patient outcomes followed up afterwards [s1].
What was tested
EyeFM is a multimodal vision-language model pretrained on 14.5 million ocular images across five imaging modalities, paired with clinical texts drawn from global, multiethnic datasets [s1]. The researchers describe it as an eyecare "copilot" — an assistant to a clinician rather than an autonomous diagnostic system.
The evaluation had several layers. Retrospective validations came first, then a multicountry efficacy validation involving 44 ophthalmologists across North America, Europe, Asia and Africa in primary and specialty care settings [s1]. The randomised trial was the last and most demanding stage: a parallel, single-centre, double-masked study of EyeFM as a copilot in retinal disease screening among a high-risk population in China [s1].
A total of 668 participants — mean age 57.5 years, 79.5% male — were randomised to 16 ophthalmologists, equally allocated between an intervention group where the clinician had the EyeFM copilot and a control group receiving standard care [s1].
The results
On the primary endpoint, ophthalmologists using EyeFM achieved a higher correct diagnostic rate than those without it: 92.2% versus 75.4% (P < 0.001) [s1]. The referral rate followed the same pattern, 92.2% versus 80.5% (P < 0.001) [s1].
A secondary outcome measured the standardisation of clinical reports, which the authors report as improved, with medians of 33 versus 37 (P < 0.001) [s1]. The abstract does not state which median belongs to which group or which direction on that scale is better, so the size and direction of that particular gain cannot be read off the summary.
The findings that arguably matter more came at follow-up, and they concern patients rather than clinicians. Participant satisfaction with the screening was similar between the two groups [s1]. But the intervention group showed higher compliance with self-management advice (70.1% versus 49.1%, P < 0.001) and with referral suggestions (33.7% versus 20.2%, P < 0.001) [s1].
That second pair of numbers is the interesting one. The AI was assisting the clinician, not speaking to the patient — yet patients in that arm were more likely to act on what they were told. The trial reports the association; it does not establish the mechanism, and the plausible explanations (better-targeted advice, more standardised reports, clinician confidence) are not distinguished by this design.
The referral compliance figures also deserve a second look in absolute terms: 33.7% in the intervention arm means about two thirds of patients advised to seek further care did not, even in the better-performing group [s1].
What the trial does not settle
It was single-centre [s1]. A screening programme in one high-risk Chinese population, run by 16 ophthalmologists at one site, is a narrow base from which to generalise, and the multicountry efficacy validation that preceded it was not itself randomised [s1].
The 79.5% male participant composition is unusual for a general screening population and unexplained by the abstract [s1], which limits what can be said about how results would transfer.
And the comparison is copilot-assisted clinicians against unassisted clinicians in a trial setting, where everyone knows a study is happening. Whether a 17-percentage-point diagnostic gap persists in routine practice, over months, once novelty and study attention fade, is the question a single trial cannot answer.
Diagnostic accuracy is also not the same as patient benefit. Higher referral rates mean more people entering the specialist system; whether that yields better vision outcomes or mostly generates workload depends on downstream capacity, which the trial did not measure.
The economics question, asked the same week
A systematic review published in npj Digital Medicine on 26 August examined the cost-effectiveness, utility and budget impact of clinical AI interventions across healthcare settings [s2]. It found 19 studies spanning oncology, cardiology, ophthalmology and infectious diseases [s2].
Nineteen studies is a small literature for a technology this widely deployed, and that is part of the finding.
The review reports that AI improved diagnostic accuracy, enhanced quality-adjusted life years and reduced costs, largely by minimising unnecessary procedures and optimising resource use, with several interventions achieving incremental cost-effectiveness ratios well below accepted thresholds [s2].
Then it qualifies that. Many evaluations relied on static models that may overestimate benefits by not capturing how AI systems change over time [s2]. Indirect costs, infrastructure investments and equity considerations were often underreported, which the authors say suggests reported economic benefits may be overstated [s2]. Dynamic modelling indicates sustained long-term value, but the review calls for further research incorporating comprehensive cost components and subgroup analyses [s2].
The underreported categories are the expensive ones. Infrastructure, integration, retraining and monitoring are where deployed AI actually costs money, and a cost-effectiveness analysis that omits them is measuring the model rather than the system.
Where this leaves the field
A well-conducted randomised trial showing that an AI copilot improved both clinician performance and patient follow-through is a meaningful step past the accuracy-benchmark literature [s1]. A systematic review finding only 19 economic evaluations, most of them structurally optimistic, indicates that the question of whether such systems pay for themselves is nowhere near settled [s2].
Both are true at once, and health systems buying these tools are making the second decision, not the first.
Sources
- An eyecare foundation model for clinical assistance: a randomized controlled trial — Nature Medicine, 2025-08-28
- Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare — npj Digital Medicine, 2025-08-26
Sources
- An eyecare foundation model for clinical assistance: a randomized controlled trial — Nature Medicine , August 28, 2025
- Systematic review of cost effectiveness and budget impact of artificial intelligence in healthcare — npj Digital Medicine , August 26, 2025
An ECG 'foundation model' matched rivals using a fraction of the labelled data
Trained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.
An AI ECG model found the one in ten older patients for whom AF screening paid off
A secondary analysis of the VITAL-AF trial reports a screening benefit only in the top risk decile — with a confidence interval whose lower bound sits at 0.01.
Medical imaging foundation models leak enough signal to re-identify patients
Off-the-shelf features from two published models matched retinal scans to the right patient 78% to 86% of the time. Fine-tuned for the task, one model reached 99.5% on OCT.
An AI reads structural heart disease off an ECG better than cardiologists can
EchoNext scored 77.3% accuracy on a 150-ECG set where 13 cardiologists averaged 64.0%. Its authors released the model weights and a 100,000-ECG labelled dataset alongside the paper.