ANALYSIS

Two June studies test AI in radiology — one for speed, one for accuracy

A draft-reporting model cut documentation time 15.5% across 23,960 radiographs without changing report quality. A separate reader study found AI assistance raised prostate MRI accuracy by 3.3 percentage points.

Most published evaluations of medical artificial intelligence measure whether a model can do something on a stored dataset. Two studies published in JAMA Network Open this month measure something narrower and more useful: what happens when radiologists actually use one.

They test different things. One asks whether AI drafting makes reporting faster without making it worse. The other asks whether AI assistance makes readers more accurate. The answers point the same way, and both are modest.

Faster reports, unchanged quality

The first study, published on 5 June, was a prospective cohort study at a tertiary care academic health system, running from 15 November 2023 to 24 April 2024 [s1]. A workflow-integrated generative model produced draft radiological reports for plain radiographs, and the analysis compared radiographs documented with model assistance against a baseline set documented without it, matched by study type — chest or non-chest [s1].

Across 23,960 radiographs — 11,980 with model use and 11,980 without — interpretations with model assistance took a mean of 159.8 seconds (SE 27.0) against 189.2 seconds (SE 36.2) without (P=0.02), a 15.5% increase in documentation efficiency [s1].

The quality question was addressed by peer review of 800 studies, which found no difference in clinical accuracy (χ²=0.68; P=0.41) or textual quality (χ²=3.62; P=0.06) between model-assisted and non-model interpretations [s1]. The textual-quality P value of 0.06 is worth noticing rather than rounding away: it is a non-significant result close to the conventional threshold, and the direction it hints at is not stated in the abstract.

A third component tested the model as a safety net. Screening 97,651 studies, it flagged those containing a clinically significant, unexpected pneumothorax with 72.7% sensitivity and 99.9% specificity [s1]. Very high specificity across nearly 100,000 studies is what makes such a flag usable without swamping a department in false alarms; 72.7% sensitivity means it would miss roughly one in four.

This was a cohort study, not a randomised trial. Radiographs were not randomly assigned to model or no-model use, and the authors describe the finding as an association rather than a causal effect of the model on documentation time [s1].

More accurate reads, by a small margin

The second study, published on 13 June, was a diagnostic study conducted between March and July 2024 using the AI system developed within the international Prostate Imaging-Cancer AI (PI-CAI) Consortium [s2].

Sixty-one readers — 34 experts and 27 non-experts — from 53 centres across 17 countries assessed biparametric prostate MRI examinations both with and without AI assistance, providing PI-RADS annotations from 3 to 5 and patient-level suspicion scores from 0 to 100 [s2]. The AI system was recalibrated on 420 Dutch examinations to generate lesion-detection maps with AI scores from 1 to 10; the remaining 360 examinations, from three Dutch centres and one Norwegian centre, formed the observer study [s2].

Among those 360 men (median age 65 years, IQR 62-70), 122 (34%) had clinically significant prostate cancer, defined by histopathology, with absence determined by three or more years of follow-up [s2].

AI assistance raised the area under the receiver operating characteristic curve by 3.3 percentage points (95% CI 1.8-4.9; P<0.001), from 0.882 (95% CI 0.854-0.910) unassisted to 0.916 (95% CI 0.893-0.938) assisted [s2]. At a PI-RADS threshold of 3 or more, sensitivity improved by 2.5 points (95% CI 1.1-3.9; P<0.001), from 94.3% to 96.8%, and specificity by 3.4 points (95% CI 0.8-6.0; P=0.01), from 46.7% to 50.1% [s2].

The specificity numbers are the interesting part. Even with AI assistance, roughly half of the men without clinically significant cancer were still scored at PI-RADS 3 or above [s2]. A 3.4-point improvement is real, and it is nowhere near solving the false-positive problem that drives unnecessary biopsy.

Secondary analyses found a greater benefit of AI assistance for non-expert readers [s2] — a recurring pattern in reader studies, and the one with the clearest implication for where such tools might matter most.

What neither study establishes

Both are observer or workflow studies, and both authors say so. The prostate study calls for further research into generalisation and into workflow effects in prospective settings [s2]. The reporting study describes its finding as suggesting potential for radiologist and generative-AI collaboration [s1].

Neither reports a patient outcome. Nobody in either study was diagnosed earlier, treated differently or survived longer because of the model. What was measured was documentation time, peer-reviewed report quality, and area under a curve.

That is not a dismissal. Those are the endpoints that determine whether a tool is adoptable at all, and they are measured here in real clinical settings across large samples rather than in retrospective benchmark exercises. But the gap between "readers score better with this" and "patients do better with this" remains unbridged, and reader studies cannot bridge it.

What to watch

Whether either system is evaluated in a randomised design with a downstream clinical endpoint — biopsy rates, cancer detection at follow-up, time to intervention — rather than reader performance or documentation time [s1][s2].

Sources

Sources

  1. Efficiency and Quality of Generative AI-Assisted Radiograph ReportingJAMA Network Open , June 5, 2025
  2. AI-Assisted vs Unassisted Identification of Prostate Cancer in Magnetic Resonance ImagesJAMA Network Open , June 13, 2025

More on

Related coverage