Doctor plus AI is meant to beat either alone. A new meta-analysis cannot show it does
Ten studies, wide confidence intervals, and factual error rates of 26 to 36 percent in AI-drafted documentation. The review's own conclusion is that the evidence is preliminary and highly uncertain.
The governing assumption behind almost every clinical deployment of a large language model is that the combination — clinician plus model — will outperform the clinician working alone. A systematic review and meta-analysis published in npj Digital Medicine on 28 January tested that assumption against the peer-reviewed literature and found the evidence thinner than the deployment curve implies [s1].
Following PRISMA 2020 and registered on PROSPERO, the authors searched four databases through 28 June 2025 for studies comparing human-plus-LLM workflows against human-only workflows [s1]. Ten peer-reviewed studies met the eligibility criteria, with three preprints used only in sensitivity analyses [s1]. That is the size of the entire comparative evidence base.
What the pooled estimates show
On diagnostic and interpretation accuracy, only two studies could be pooled. The point estimate favoured the combination — a risk ratio of 1.59 — but the confidence interval ran from 0.08 to 32.74, and the 95% prediction interval crossed the null [s1]. An interval that wide is not a weak positive finding; it is an absence of information.
On composite diagnostic and management scores, again from two studies, the combination was statistically better by a mean difference of 4.88 percentage points (95% CI +0.65 to +9.12) [s1]. But the prediction interval — the range within which a new study's result would be expected to fall — ran from −31.65 to +41.42 [s1]. The authors' own reading is that this indicates high real-world uncertainty [s1].
On time efficiency, pooling three studies, there was no overall difference: a mean difference of 0.4 minutes (95% CI −4.18 to +4.97), with substantial heterogeneity (I² = 70.1%) [s1].
Documentation quality improved. But factual error rates in the reviewed work remained in the range of roughly 26% to 36%, which the authors say undermines the quality gains [s1].
The finding most likely to unsettle procurement decisions concerns three-arm studies — those comparing human-only, human-plus-AI and AI-only. In those settings, human-plus-AI did not universally outperform AI alone [s1]. The collaborative configuration that health systems are building toward is not consistently the best-performing of the three in the studies that bothered to measure all three.
The review's stated conclusion is that the evidence remains preliminary, highly uncertain and context-dependent, and it recommends preregistered, pragmatic, multicentre trials embedded in real workflows, with harmonised core outcomes that prioritise safety and error metrics, and interfaces that surface uncertainty and support verification [s1].
What the underlying studies tend to measure
A simulation study published in JCO Clinical Cancer Informatics on the same day illustrates both the promise and the measurement problem [s2]. Twenty-six oncologists from the United Kingdom, United States, Spain and Singapore reviewed synthetic breast cancer cases and produced tumour board summaries twice: once using an LLM-enabled clinical decision support platform that supplied editable generated summaries, and once using a simulated electronic health record requiring manual composition [s2].
The platform cut median summary completion time from 8 minutes 47 seconds to 6 minutes 55 seconds [s2]. Completeness scores improved; correctness and conciseness were similar between conditions [s2]. Eighty-seven per cent of participants said they would recommend the platform and 96% anticipated time savings, with a System Usability Scale score of 65.7 [s2]. Perceived cognitive load was lower with the platform, but the difference was not statistically significant [s2].
Read carefully, this is a study of speed and completeness on synthetic cases, with correctness unchanged — which is close to what the meta-analysis describes as the general pattern: efficiency and perceived quality move; accuracy and safety endpoints largely do not, or are not measured at all [s1].
Clinicians are not the obstacle people assume
A qualitative systematic review and meta-synthesis published in JMIR AI on 5 February pooled 13 studies from six countries, covering qualitative data from 238 primary care physicians, nurses, physiotherapists and other professionals providing direct patient care [s3]. Eight descriptive themes were synthesised into three analytical themes: the human–machine relationship, the technologically enhanced clinic, and the societal impact of AI, the last covering data privacy, medicolegal liability and bias [s3]. Confidence in the findings, assessed with GRADE-CERQual, was rated high for 15 findings, moderate for five and low for one [s3].
The synthesis describes clinicians as viewing AI as a technology that can both enhance and complicate primary care, with integration requiring attention to ethical implications, technical reliability and the maintenance of human oversight [s3]. The authors note that interpretation is constrained by heterogeneity in qualitative methods and by the diversity of technologies studied [s3].
The limits worth stating plainly
Two of the meta-analysis's key pooled estimates rest on two studies each [s1]. Heterogeneity was high where it could be measured [s1]. The search closed in June 2025, so tools released or revised since then are outside it [s1]. And a systematic review can only summarise what was studied — if the field has largely measured minutes saved and satisfaction rather than diagnostic error or patient harm, the pooled evidence will reflect that gap rather than fill it.
What to watch
Whether any of the pragmatic, multicentre trials the review calls for are registered, and whether they adopt safety and error rates as primary outcomes rather than efficiency [s1]. And whether three-arm designs become standard — because until they are routine, the question of whether the human in human-plus-AI is adding accuracy or only adding time will keep going unanswered [s1].
Sources
- [s1] Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis. npj Digital Medicine, published online 28 January 2026. https://doi.org/10.1038/s41746-026-02382-2
- [s2] Simulation-Based Evaluation of a Large Language Model-Enabled Clinical Decision Support Platform in Oncology. JCO Clinical Cancer Informatics, 28 January 2026. https://doi.org/10.1200/cci-25-00244
- [s3] Exploring Clinician Perspectives on Artificial Intelligence in Primary Care: Qualitative Systematic Review and Meta-Synthesis. JMIR AI, published online 5 February 2026. https://doi.org/10.2196/72210
Sources
- Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis — npj Digital Medicine , January 28, 2026
- Simulation-Based Evaluation of a Large Language Model-Enabled Clinical Decision Support Platform in Oncology — JCO Clinical Cancer Informatics , January 28, 2026
- Exploring Clinician Perspectives on Artificial Intelligence in Primary Care: Qualitative Systematic Review and Meta-Synthesis — JMIR AI , February 5, 2026
A computer that grades each colonoscopy raised how often endoscopists found adenomas
A Danish stepped-wedge trial gave endoscopists automated feedback on their technique after every procedure. Adenoma detection rose from 43.4% to 48.6% — a different tool from real-time polyp AI.
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
Most AI models that predict who will skip their medicines aren't ready for the clinic
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.
Models beat German medical students on text — and fell apart on the picture questions
Across 24 official German licensing exams, the best model answered 99.31% of first-exam items correctly. On items containing an image, the error rate rose several-fold, against 1.24x for students.