Fracture-spotting AI is accurate. A children's ED trial asked if that changes care
Commercial software reads an X-ray for breaks about as well as a radiologist. A prospective paediatric emergency study found that accuracy did not translate into fewer missed-fracture recalls.
| Group | Value (%) |
|---|---|
| Without AI support | 8.6 |
| With AI support | 5.7 |
Fracture-detection AI is software that reads a plain X-ray and flags the breaks, drawing a box around a suspected fracture to prompt the clinician who missed it. On the narrow question of whether it can do that, the evidence is settled: pooled across dozens of studies, these systems match human readers for accuracy [s2]. On the question that matters more — whether adding one to a busy emergency department actually changes what happens to patients — a prospective study in a children's ED has delivered a deflating answer [s1].
The accuracy is real
A systematic review and meta-analysis in Radiology, drawing on 42 studies and more than 55,000 images, put AI's pooled sensitivity for detecting fractures at 92% (95% CI 88–93) with specificity of 91% (95% CI 88–93) on internal test sets — statistically indistinguishable from the clinicians it was compared against, who scored 91% sensitivity and 92% specificity [s2]. On the harder external validation sets, AI held at 91% sensitivity and 91% specificity, a little behind clinicians at 94% on both, but again with no significant difference [s2].
Fractures are a good target for this kind of tool because they are common, they are missed often enough to matter — a subtle break on an odd projection, read at 3am — and a miss has a clear cost. So the case for putting a second reader in the room is intuitive. The test is whether the intuition survives contact with a real department.
What the children's ED study did
Researchers at a German tertiary hospital ran a prospective study during out-of-hours care between April and September 2025 [s1]. Children and adolescents aged 2 to 18 having a limb X-ray were eligible; on every second day, the treating physician had access to automated fracture detection from a commercial system (TechCare Kids, made by Milvue), and on the alternate days they did not [s1]. Of 1,515 patients screened, 667 were enrolled, median age 11, and fractures were genuinely present in 296 of them — 44.3% [s1].
The AI itself performed as advertised: its accuracy in this real-world stream was 95.1% [s1]. The primary endpoint, though, was not accuracy but consequence — the rate at which the initial read had to be revised, bringing the child back the next day.
The result the accuracy did not buy
Diagnostic revisions leading to a recall occurred in 8.6% of cases without AI and 5.7% with it — a reduction, but one whose confidence interval comfortably includes no effect at all (risk ratio 0.66, 95% CI 0.37–1.19) [s1]. Downstream changes to treatment were rare either way, 2.0% without AI versus 0.4% with, and the difference was not significant (p=0.10) [s1]. Nor did AI move the secondary measures: the need for a senior consultation (14.3% versus 12.5%), the physicians' own confidence in their read, or length of stay in the department (2.3 versus 2.4 hours) all looked essentially the same [s1].
The authors' conclusion is unusually candid for a technology paper. Given how small the effect on recalls was, they wrote, it is questionable whether the expense of a larger, adequately powered confirmatory trial would be justified by the incremental benefit on offer [s1].
Why accuracy and benefit come apart
The gap is not a paradox; it is the normal shape of a good diagnostic tool meeting an already-capable system. If clinicians in a tertiary paediatric ED are reading limb films well to begin with, an accurate AI has little headroom to prevent errors that were not being made — and the fractures it catches may be the small, stable ones whose management does not change. A tool can be 95% accurate and still, at the margin, alter almost nothing that a patient would feel.
That is the recurring lesson of the site's coverage of clinical AI: a device clearance certifies that software performs a task, not that deploying it improves care, and the two are separated by exactly the outcome evidence that is usually missing. It is why an accurate stroke-imaging AI still has to prove it changes treatment rates, and why regulators have been pressed on the thinness of real-world testing behind cleared radiology tools. Detection is the easy half of the claim.
What to watch
Where fracture AI is tested next. The setting that could show a benefit is not a well-staffed teaching hospital but the one this study was not run in — an overnight or remote department without a radiologist on hand, where a second read is otherwise unavailable. A single prospective study in one paediatric ED does not close the question; it reframes it, from "is the software accurate" to "where, if anywhere, does that accuracy earn its keep."
This article is informational and is not medical advice.
Sources
- [s1] Deffaa OJ, Pape J, Schlösser D, et al. "Artificial intelligence for pediatric fracture detection: impact on diagnostic revisions and patient recall rates in a tertiary emergency setting." BMC Emergency Medicine, 26(1):204, published online 29 July 2026. https://doi.org/10.1186/s12873-026-01697-3
- [s2] Kuo RYL, Harrison C, Curran T-A, et al. "Artificial Intelligence in Fracture Detection: A Systematic Review and Meta-Analysis." Radiology, 304(1):50–62, published online 29 March 2022. https://doi.org/10.1148/radiol.211785
Sources
- Artificial intelligence for pediatric fracture detection: impact on diagnostic revisions and patient recall rates in a tertiary emergency setting — BMC Emergency Medicine , July 29, 2026
- Artificial Intelligence in Fracture Detection: A Systematic Review and Meta-Analysis — Radiology , March 29, 2022
More on
AI for pulmonary embolism on CT buys speed, not accuracy — three 2026 studies
Where AI clot-detection was tested head-to-head, it still missed more emboli than radiology trainees, especially small peripheral ones. Its clearest win was cutting time to diagnosis.
AI can flag pancreatic cancer on ordinary CT scans, but only in retrospective tests
A deep-learning model reads the pancreas on non-contrast CT taken for other reasons. It reached very high accuracy in large tests — none of them a prospective screening trial.
AI reads a fetal heart trace better. No trial shows that helps the baby.
Software that interprets cardiotocography in labour lifts clinicians' accuracy in reader studies. The largest randomised trial of computerised interpretation, in 47,062 women, found no benefit for mothers or babies.
A blood-pressure algorithm predicts surgical hypotension. A simple alarm matched it.
The Hypotension Prediction Index warns anaesthetists before blood pressure falls. Three 2026 trials find a plain mean-pressure alarm does the same job, and the algorithm's benefit tracks how hard clinicians treat.