An AI ECG model found the one in ten older patients for whom AF screening paid off
A secondary analysis of the VITAL-AF trial reports a screening benefit only in the top risk decile — with a confidence interval whose lower bound sits at 0.01.
Screening everyone over 65 for atrial fibrillation has repeatedly produced disappointing returns. An analysis of the VITAL-AF trial, published in the Journal of the American College of Cardiology, tests a different premise: that the problem is not screening but who gets screened, and that an AI model reading a routine ECG can sort the two groups apart [s1].
The trial and the question
VITAL-AF was a cluster-randomised trial of patients aged 65 and over treated at one of 16 primary care practices affiliated with Massachusetts General Hospital [s1]. Patients randomised to a screening practice were screened using a single-lead ECG [s1]. The trial's own starting point is stated in the paper: current screening approaches using the guideline age threshold of 65 or over have shown limited yield [s1].
The new analysis takes participants who did not already have AF and who had at least one 12-lead ECG in the three years before enrolment, and applies three risk models — all of them validated outside VITAL-AF — to see whether the screening effect concentrates in higher-risk patients [s1].
The three models were the CHARGE-AF clinical score, an AI model using a 12-lead ECG alone (ECG-AI), and a combination of the two (CH-AI) [s1]. Of 30,630 VITAL-AF participants without prevalent AF, 16,937 had the pretrial ECG and clinical data needed [s1].
How well the models sorted people
All three discriminated two-year AF risk. Areas under the receiver-operating characteristic curve were 0.711 (95% CI 0.671-0.749) for CHARGE-AF, 0.784 (95% CI 0.743-0.819) for ECG-AI, and 0.788 (95% CI 0.754-0.824) for the combined CH-AI model [s1].
The average precision figures are the more informative ones, and they are the numbers that get lost in coverage of this kind of result: 0.0952 (95% CI 0.0836-0.112) for CHARGE-AF, 0.132 (95% CI 0.113-0.157) for ECG-AI, and 0.133 (95% CI 0.117-0.159) for CH-AI [s1]. Average precision reflects performance against a rare outcome in a way AUROC does not. The AI-based models roughly outperform the clinical score on this measure — and all three remain low in absolute terms, because incident AF over two years is uncommon even in this age group.
Where screening actually helped
The analysis defines the screening effect as the difference in two-year incident AF diagnosis rate between screening and control practices, examined across deciles of predicted risk [s1].
A screening effect appeared in the top decile of the CH-AI model. In that decile, AF was diagnosed at 10.07 per 100 person-years in the screening arm (95% CI 8.28-11.87) against 7.76 in the control arm (95% CI 6.30-9.21), with P < 0.05 [s1]. The difference is 2.32 per 100 person-years, with a 95% confidence interval of 0.01 to 4.63, corresponding to a number needed to screen of 43 per year [s1].
That interval deserves to be read out loud. Its lower bound is 0.01 per 100 person-years — a value functionally indistinguishable from no effect. The result clears statistical significance by the narrowest possible margin. It is a signal, not a settled quantity, and the number needed to screen of 43 is a point estimate sitting on top of that uncertainty.
The trade-off the authors name
The paper's own conclusion is careful. It says the findings suggest a trade-off between increasing AF screening efficiency and decreasing population coverage — that is, restricting the pool of people screened [s1]. Concentrating screening on the top decile means nine in ten patients who would have been screened under the age-based rule are no longer screened, and some of them would have had AF found.
The authors state that future studies are needed to determine whether a risk-based approach is optimal, or whether additional clinical and systems-level factors such as access and health care system engagement could further refine screening strategies [s1].
What this analysis is and is not
It is a secondary analysis of a completed randomised trial, using risk models applied retrospectively to stored ECGs. It is not a prospective test of a risk-guided screening programme. Nobody in VITAL-AF was selected for screening on the basis of an AI score; the score was calculated afterwards and used to stratify people who had already been randomised by practice.
The distinction matters because a deployed risk-guided programme would face problems this analysis cannot see: patients without a recent 12-lead ECG on file, clinicians deciding whether to act on a score, and the behaviour of a model on a population that differs from the one it was derived on. Of VITAL-AF's participants without prevalent AF, 16,937 of 30,630 had the prior ECG and clinical data this approach requires [s1]. In a real programme, the roughly 45% who did not are not a rounding error; they are the patients the strategy has no answer for.
What to watch
Whether a prospective trial randomises patients to risk-guided versus age-based AF screening rather than reconstructing the comparison after the fact. Whether the effect in the top risk decile replicates in a population outside a single academic health system's primary care network. And whether any AI-ECG risk model is evaluated on the outcome that motivates AF screening in the first place — stroke — rather than on AF diagnosis, which is the intermediate step.
Sources
- [s1] Risk-Guided Atrial Fibrillation Screening With Artificial Intelligence-Enabled Electrocardiogram Models: A VITAL-AF Trial Analysis. Journal of the American College of Cardiology, April 2026 (available online 14 April 2026). ClinicalTrials.gov NCT03515057. https://doi.org/10.1016/j.jacc.2026.01.087
Sources
- Risk-Guided Atrial Fibrillation Screening With Artificial Intelligence-Enabled Electrocardiogram Models: A VITAL-AF Trial Analysis — Journal of the American College of Cardiology , April 14, 2026
More on
An ECG 'foundation model' matched rivals using a fraction of the labelled data
Trained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.
Ablation for atrial fibrillation beat a sham procedure by 2.6 points, which is nothing
The first double-blind sham-controlled trial of pulmonary vein isolation found quality of life improved in both arms. Most of the benefit patients feel appears not to come from the ablation.
Catheter ablation after a stroke didn't cut second-stroke risk, a Japanese trial found
The STABLED trial randomized 331 atrial fibrillation patients with a recent stroke to ablation plus anticoagulation or anticoagulation alone. The two groups ended up nearly identical.
A trial finally tested blood thinners in the atrial fibrillation grey zone
Guidelines have hedged on anticoagulation for people with one stroke risk factor. SINGLE-AF randomised 1,803 patients in South Korea, and the events it counted were very few in both arms.