WHAT THE STUDY ACTUALLY SAYS

AI mammography held interval-cancer rates and raised sensitivity in the MASAI trial

The first randomised trial to report interval cancers found AI non-inferior to double reading, with higher sensitivity at equal specificity. Across three European programmes the extra detection is about one per 1000.

Screening sensitivity: AI-supported vs standard double readingAI-supported screening: 80.5%; Standard double reading: 73.8%0%45%90%AI-supported screening80.5%Standard double reading73.8%
Screening sensitivity: AI-supported vs standard double reading
GroupValue (%)
AI-supported screening80.5 (76.4 to 84.2)
Standard double reading73.8 (68.9 to 78.3)
Screening sensitivity: AI-supported vs standard double reading MASAI, a Swedish population-based randomised trial; 105,934 women allocated 1:1. Source: The Lancet

The first randomised trial to report the effect of AI-supported mammography on interval cancers found the rate no worse than standard double reading, while sensitivity was higher at the same specificity [s1]. That is a genuinely positive result for a technology usually judged on retrospective accuracy alone — but a separate synthesis of three European programmes puts the extra cancers found at roughly one per 1000 women screened, which is the number a health system actually has to weigh [s2].

What MASAI measured

MASAI is a Swedish population-based, randomised, single-blinded, non-inferiority trial run through a routine screening programme [s1]. Between 12 April 2021 and 7 December 2022, 105,934 women were allocated 1:1 to AI-supported screening or standard double reading without AI, with 19 excluded from analysis [s1]. In the intervention group the AI both triaged each examination to single or double radiologist reading and flagged suspicious areas as detection support [s1].

The primary outcome here is the interval cancer rate — cancers that surface between screening rounds or within two years of a clear screen, the ones a screening test missed [s1]. Earlier MASAI reports had shown gains in cancer detection and in workload; whether AI would also reduce the cancers screening fails to catch was, as the authors put it, unknown [s1]. The analysis was powered against a 20% non-inferiority margin [s1].

The results

Interval cancer rates were 1.55 per 1000 participants (95% CI 1.23–1.92) in the AI group and 1.76 (1.42–2.15) in the control group — a proportion ratio of 0.88 (95% CI 0.65–1.18; p=0.41), within the non-inferiority margin [s1]. The point estimate favours AI, but the confidence interval crosses 1, so this is a demonstration that AI did not make interval cancers worse, not proof that it reduced them.

The descriptive breakdown is more suggestive. The AI group had fewer interval cancers that were invasive (75 vs 89), fewer that were T2 or larger (38 vs 48), and fewer of the more aggressive non-luminal A subtype (43 vs 59) [s1]. These are counts, not tested differences, but they point the same direction: the cancers that slipped through in the AI arm tended to be less advanced.

On the secondary outcome that most directly measures reading performance, sensitivity was 80.5% (95% CI 76.4–84.2) with AI versus 73.8% (68.9–78.3) without it (p=0.031), a difference consistent across age and breast density and present for invasive but not in-situ cancer [s1]. Specificity was identical at 98.5% in both groups (p=0.88) [s1]. Higher sensitivity with no loss of specificity is the combination screening programmes want and rarely get; usually catching more means recalling more healthy women.

The conflicts worth naming

The trial was publicly funded, by the Swedish Cancer Society, the regional cancer centres and Swedish governmental research funding [s1]. The disclosures are still worth reading: the senior author reports advisory board membership with Siemens Healthineers and is co-founder of an AI point-of-care ultrasound company, and a co-author's institution, BreastScreen Norway, holds a research agreement with ScreenPoint Medical, whose software is used in this field [s1]. None of that impeaches a publicly funded randomised trial, but it is the kind of provenance a reader is entitled to see stated rather than buried.

The number a system has to weigh

A meta-analysis published in Clinical Imaging in August 2026 pooled the three large prospective evaluations embedded in European-style screening — MASAI, the paired-reader ScreenTrustCAD study, and the nationwide German PRAIM implementation — across 597,419 examinations [s2]. Its estimate of the cancer-detection benefit is deliberately modest: a pooled absolute increase of 0.9 additional cancers per 1000 examinations (95% CI −0.0 to +1.8), with low heterogeneity [s2].

Crucially, that gain came without a consistent rise in recall: the pooled recall difference was −0.6 per 1000 (95% CI −3.1 to +2.1) [s2]. The efficiency signal was large — MASAI cut total readings by 44.3%, and in PRAIM a programme-level safety net recovered 204 cancers that would otherwise have been missed [s2]. The authors' framing is that AI works best as a complementary reader inside these workflows, conditional on explicit quality assurance and ongoing monitoring of interval cancers and stage distribution [s2].

That "about one per 1000" is easy to read as underwhelming and easy to read as substantial; both are wrong without the denominator. Applied across a national programme screening millions, one extra cancer per 1000 is a large number of women; applied to an individual visit, it is a small shift in the odds. The honest statement is that the benefit is real, replicated, and modest per examination — and that the workload reduction may be the more immediately decisive finding for stretched services.

What to watch

MASAI settles that AI-supported reading does not increase the cancers screening misses, at least over one interval, in one Swedish programme [s1]. What it cannot settle is the long-run stage-distribution and mortality question that only years of follow-up answer, or whether the same result holds with different software, different populations and different reader cultures — the reasons the pooled analysis still reports meaningful heterogeneity on recall [s2]. The regulatory literature has repeatedly found AI imaging devices cleared on accuracy metrics without evidence they change patient outcomes; MASAI is a step past that, not the end of it. For the wider debate on that outcomes gap, see our coverage of the FDA's AI radiology testing gap, two 2026 radiology-assist studies, and risk-based breast screening in the WISDOM trial.

Sources

Sources

  1. Interval cancer, sensitivity, and specificity comparing AI-supported mammography screening with standard double reading without AI in the MASAI study: a randomised, controlled, non-inferiority, single-blinded, population-based, screening-accuracy trialThe Lancet , January 31, 2026
  2. Artificial intelligence–supported double reading in European population breast cancer screening: A systematic review and meta-analysis of prospective programsClinical Imaging , August 11, 2026

More on

Related coverage