WHAT THE STUDY ACTUALLY SAYS

AI reads a fetal heart trace better. No trial shows that helps the baby.

Software that interprets cardiotocography in labour lifts clinicians' accuracy in reader studies. The largest randomised trial of computerised interpretation, in 47,062 women, found no benefit for mothers or babies.

Clinicians predicting fetal acidaemia in a reader trial, unaided vs computer-assistedOverall success, unaided: 54%; Overall success, with cCTG: 61.4%; Sensitivity, unaided: 49.3%; Sensitivity, with cCTG: 61.7%; Specificity, unaided: 58.7%; Specificity, with cCTG: 61.2%0%35%70%Overall success, unaided54%Overall success, with cCTG61.4%Sensitivity, unaided49.3%Sensitivity, with cCTG61.7%Specificity, unaided58.7%Specificity, with cCTG61.2%
Clinicians predicting fetal acidaemia in a reader trial, unaided vs computer-assisted
GroupValue (%)
Overall success, unaided54
Overall success, with cCTG61.4
Sensitivity, unaided49.3
Sensitivity, with cCTG61.7
Specificity, unaided58.7
Specificity, with cCTG61.2
Clinicians predicting fetal acidaemia in a reader trial, unaided vs computer-assisted 211 clinicians assessed 100 traces, half from births with cord-blood pH below 7.15, with and without computerised CTG (cCTG) assistance. Source: npj Digital Medicine

Cardiotocography (CTG) is the continuous trace of the fetal heart rate and the mother's contractions that most women in hospital are monitored with during labour, and software has long promised to read that trace more reliably than a tired clinician at 3am. The honest answer from the trial evidence is split: computer assistance does make clinicians modestly better at spotting a distressed baby in reader studies [s3], but the largest randomised trial of computerised interpretation — 47,062 women — found it changed nothing that matters for mothers or babies [s1].

What the technology is trying to fix

Reading a CTG is a pattern-recognition task done under time pressure, and clinicians disagree with each other a lot: interpretation "is subject to high interobserver variability, limiting its performance for predicting perinatal acidaemia," the low cord-blood pH that signals a baby was starved of oxygen [s3]. That variability is exactly what an algorithm is meant to remove, and computerised interpretation of the fetal heart rate has been attempted for decades. The problem sits one layer down, in the test itself.

What continuous monitoring alone achieves

Before asking whether software reads CTG well, it is worth asking whether continuous CTG helps at all. A Cochrane review of 13 trials in more than 37,000 women compared continuous CTG with a midwife intermittently listening to the fetal heart [s2]. Continuous monitoring halved neonatal seizures (risk ratio 0.50, 95% CI 0.31 to 0.80) — a real benefit — but produced no reduction in perinatal death (RR 0.86, 95% CI 0.59 to 1.23) and no difference in cerebral palsy (RR 1.75, 95% CI 0.84 to 3.63) [s2]. It also increased caesarean sections (RR 1.63, 95% CI 1.29 to 2.07) and instrumental vaginal births (RR 1.15, 95% CI 1.01 to 1.33) [s2]. The machine that generates the trace, in other words, buys one narrow benefit at the cost of more surgery — because it flags far more alarm than there is genuine distress.

The landmark test of computerised reading

If human variability were the main problem, adding a computer should help. The INFANT trial set out to prove it. Between 6 January 2010 and 31 August 2013 it randomised 47,062 women having continuous electronic monitoring to either decision-support software or none (23,515 versus 23,547), analysing 46,042 [s1]. The result was flat. Poor neonatal outcome occurred in 172 of 22,987 babies (0.7%) with the software and 171 of 23,055 (0.7%) without, an adjusted risk ratio of 1.01 (95% CI 0.82 to 1.25), and there were no significant differences in developmental assessment at 2 years [s1]. The authors' conclusion was blunt: computerised interpretation of cardiotocographs "does not improve clinical outcomes for mothers or babies" [s1].

The new AI wave, and where it genuinely helps

Machine learning has since revisited the problem, and here the picture is more encouraging — as long as you are careful about what is being measured. A 2026 randomised multi-reader study in npj Digital Medicine had 211 clinicians from 23 countries each assess 100 CTG recordings, half of them from births with a cord pH below 7.15, with or without computerised CTG (cCTG) assistance [s3]. The assist helped: the overall success rate rose from 54.0% to 61.4% (p < 0.01) and sensitivity from 49.3% to 61.7% (p < 0.01), while specificity did not change significantly (58.7% versus 61.2%, p = 0.14) [s3]. In cases where clinician and model disagreed, the model was right 67.5% of the time [s3]. That is a real improvement in how accurately doctors read a trace — but it is a reader study, not a trial of what happens to babies.

The derivation studies keep arriving with headline accuracy. A 2026 random-forest model built to predict an umbilical-artery pH below 7.20 reported 91.0% accuracy (area under the curve 0.95) from heart-rate features alone, and 93.0% with added maternal and perinatal factors [s4]. Its own authors supply the caveat: it was retrospective and case-enriched, and "external validation in independent consecutive populations is required before clinical implementation" [s4].

Why accuracy on a trace need not reach the baby

The recurring gap is between classifying a trace and preventing harm. Continuous CTG's own weakness is poor specificity — it raises the alarm on many babies who are fine, which is why it drives up caesareans without cutting death or cerebral palsy [s2]. Software that sorts traces more consistently can still inherit that ceiling, because the signal it is sorting is only loosely tied to the outcome anyone cares about. INFANT tested computerised reading against that ceiling in tens of thousands of real labours and could not move it [s1]. ACOG's 2025 guideline still sets intrapartum fetal heart-rate monitoring as the framework for evaluation and management, and devotes much of its length to standardising the nomenclature and classification of patterns [s5] — a tacit acknowledgement that interpretation, not hardware, is where the uncertainty lives.

What it means, and what to watch

None of this says AI has no place in the delivery room; better, more consistent reads of an ambiguous trace are worth having [s3]. It says the specific claim being marketed — that AI improves fetal monitoring — is so far a claim about traces, not about babies. The study that would change the verdict is a randomised trial that assigns AI-guided versus standard CTG and counts neonatal acidaemia, encephalopathy, seizures and caesareans, not reader accuracy on a slideshow of recordings. Until one exists, the burden of proof sits with the algorithm, and the honest summary is that the last time anyone ran that trial at scale, the computer made no difference [s1].

Sources

Related coverage: why cleared AI devices so rarely show a patient-outcome benefit, the testing gap in FDA-cleared AI radiology tools, and an algorithm that lost to a plain blood-pressure alarm.

Sources

  1. Computerised interpretation of fetal heart rate during labour (INFANT): a randomised controlled trial — The Lancet , April 1, 2017
  2. Continuous cardiotocography (CTG) as a form of electronic fetal monitoring (EFM) for fetal assessment during labour — Cochrane Database of Systematic Reviews , February 3, 2017
  3. Randomised study of human machine collaboration for cardiotocography interpretation during labour — npj Digital Medicine , March 19, 2026
  4. Fetal Acidemia Prediction Using a Random Forest Model With Cumulative Fetal Stress Assessment During Labor — Journal of Obstetrics and Gynaecology Research , August 20, 2026
  5. ACOG Clinical Practice Guideline No. 10: Intrapartum Fetal Heart Rate Monitoring: Interpretation and Management — Obstetrics & Gynecology , October 1, 2025
Related coverage