ANALYSIS

Doctor plus AI is meant to beat either alone. A new meta-analysis cannot show it does

Ten studies, wide confidence intervals, and factual error rates of 26 to 36 percent in AI-drafted documentation. The review's own conclusion is that the evidence is preliminary and highly uncertain.

The governing assumption behind almost every clinical deployment of a large language model is that the combination — clinician plus model — will outperform the clinician working alone. A systematic review and meta-analysis published in npj Digital Medicine on 28 January tested that assumption against the peer-reviewed literature and found the evidence thinner than the deployment curve implies [s1].

Following PRISMA 2020 and registered on PROSPERO, the authors searched four databases through 28 June 2025 for studies comparing human-plus-LLM workflows against human-only workflows [s1]. Ten peer-reviewed studies met the eligibility criteria, with three preprints used only in sensitivity analyses [s1]. That is the size of the entire comparative evidence base.

What the pooled estimates show

On diagnostic and interpretation accuracy, only two studies could be pooled. The point estimate favoured the combination — a risk ratio of 1.59 — but the confidence interval ran from 0.08 to 32.74, and the 95% prediction interval crossed the null [s1]. An interval that wide is not a weak positive finding; it is an absence of information.

On composite diagnostic and management scores, again from two studies, the combination was statistically better by a mean difference of 4.88 percentage points (95% CI +0.65 to +9.12) [s1]. But the prediction interval — the range within which a new study's result would be expected to fall — ran from −31.65 to +41.42 [s1]. The authors' own reading is that this indicates high real-world uncertainty [s1].

On time efficiency, pooling three studies, there was no overall difference: a mean difference of 0.4 minutes (95% CI −4.18 to +4.97), with substantial heterogeneity (I² = 70.1%) [s1].

Documentation quality improved. But factual error rates in the reviewed work remained in the range of roughly 26% to 36%, which the authors say undermines the quality gains [s1].

The finding most likely to unsettle procurement decisions concerns three-arm studies — those comparing human-only, human-plus-AI and AI-only. In those settings, human-plus-AI did not universally outperform AI alone [s1]. The collaborative configuration that health systems are building toward is not consistently the best-performing of the three in the studies that bothered to measure all three.

The review's stated conclusion is that the evidence remains preliminary, highly uncertain and context-dependent, and it recommends preregistered, pragmatic, multicentre trials embedded in real workflows, with harmonised core outcomes that prioritise safety and error metrics, and interfaces that surface uncertainty and support verification [s1].

What the underlying studies tend to measure

A simulation study published in JCO Clinical Cancer Informatics on the same day illustrates both the promise and the measurement problem [s2]. Twenty-six oncologists from the United Kingdom, United States, Spain and Singapore reviewed synthetic breast cancer cases and produced tumour board summaries twice: once using an LLM-enabled clinical decision support platform that supplied editable generated summaries, and once using a simulated electronic health record requiring manual composition [s2].

The platform cut median summary completion time from 8 minutes 47 seconds to 6 minutes 55 seconds [s2]. Completeness scores improved; correctness and conciseness were similar between conditions [s2]. Eighty-seven per cent of participants said they would recommend the platform and 96% anticipated time savings, with a System Usability Scale score of 65.7 [s2]. Perceived cognitive load was lower with the platform, but the difference was not statistically significant [s2].

Read carefully, this is a study of speed and completeness on synthetic cases, with correctness unchanged — which is close to what the meta-analysis describes as the general pattern: efficiency and perceived quality move; accuracy and safety endpoints largely do not, or are not measured at all [s1].

Clinicians are not the obstacle people assume

A qualitative systematic review and meta-synthesis published in JMIR AI on 5 February pooled 13 studies from six countries, covering qualitative data from 238 primary care physicians, nurses, physiotherapists and other professionals providing direct patient care [s3]. Eight descriptive themes were synthesised into three analytical themes: the human–machine relationship, the technologically enhanced clinic, and the societal impact of AI, the last covering data privacy, medicolegal liability and bias [s3]. Confidence in the findings, assessed with GRADE-CERQual, was rated high for 15 findings, moderate for five and low for one [s3].

The synthesis describes clinicians as viewing AI as a technology that can both enhance and complicate primary care, with integration requiring attention to ethical implications, technical reliability and the maintenance of human oversight [s3]. The authors note that interpretation is constrained by heterogeneity in qualitative methods and by the diversity of technologies studied [s3].

The limits worth stating plainly

Two of the meta-analysis's key pooled estimates rest on two studies each [s1]. Heterogeneity was high where it could be measured [s1]. The search closed in June 2025, so tools released or revised since then are outside it [s1]. And a systematic review can only summarise what was studied — if the field has largely measured minutes saved and satisfaction rather than diagnostic error or patient harm, the pooled evidence will reflect that gap rather than fill it.

What to watch

Whether any of the pragmatic, multicentre trials the review calls for are registered, and whether they adopt safety and error rates as primary outcomes rather than efficiency [s1]. And whether three-arm designs become standard — because until they are routine, the question of whether the human in human-plus-AI is adding accuracy or only adding time will keep going unanswered [s1].

Sources

  1. [s1] Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis. npj Digital Medicine, published online 28 January 2026. https://doi.org/10.1038/s41746-026-02382-2
  2. [s2] Simulation-Based Evaluation of a Large Language Model-Enabled Clinical Decision Support Platform in Oncology. JCO Clinical Cancer Informatics, 28 January 2026. https://doi.org/10.1200/cci-25-00244
  3. [s3] Exploring Clinician Perspectives on Artificial Intelligence in Primary Care: Qualitative Systematic Review and Meta-Synthesis. JMIR AI, published online 5 February 2026. https://doi.org/10.2196/72210

Sources

  1. Human–large language model collaboration in clinical medicine: a systematic review and meta-analysisnpj Digital Medicine , January 28, 2026
  2. Simulation-Based Evaluation of a Large Language Model-Enabled Clinical Decision Support Platform in OncologyJCO Clinical Cancer Informatics , January 28, 2026
  3. Exploring Clinician Perspectives on Artificial Intelligence in Primary Care: Qualitative Systematic Review and Meta-SynthesisJMIR AI , February 5, 2026
Related coverage