Two 2026 trials put language-model decision support in front of real patients
One found doctors stopped using it as shifts got busier. The other, in 9,691 Kenyan patients, found it safe but no better than the electronic record alone. Neither supports deployment.
Most published evidence on large language models in medicine comes from benchmarks — the model sits an exam, or is scored against a set of vignettes. Two trials published in Nature Medicine in 2026 did something rarer and more informative: they put the systems into working clinics and measured what happened to patients and clinicians.
Both are negative or near-negative on their headline questions. Both are more useful than a positive benchmark.
The emergency department
The first was a DECIDE-AI stage 1 evaluation of SHAKED, a clinical decision support system built on multiple large language models, in a tertiary emergency department [s1]. Over four weeks, 1,138 patients were analysed across two parallel units — one using the system, one following routine rotations [s1].
On the quality of the system's output, the result was strong. Expert review rated 99 of 100 sampled outputs as clinically appropriate, and no adverse events were detected [s1].
On whether it changed anything, it did not. Emergency department length of stay was 4.9 hours in both wings (P = 0.99) [s1]. An intention-to-treat analysis showed a non-significant trend toward shorter consultation cycle time, at −9.4 minutes (P = 0.077) [s1].
The finding the authors lead with is neither of those. Clinical adoption of the system declined from 68% to 30% over the four weeks, driven by what they term workload-sensitive disengagement, with an odds ratio of 0.72 per shift hour (95% CI, 0.62 to 0.83) [s1].
Read that mechanism carefully. Use fell as shifts wore on — the busier the clinician, the less likely they were to consult the tool. That is the exact inverse of the value proposition. Decision support is sold as most valuable under pressure, and this system was abandoned under pressure.
Physicians did keep using it for one thing: they preferred it for radiology consultations, with an odds ratio of 2.98 (95% CI, 1.58 to 5.63) [s1]. A tool used selectively for the task clinicians find it good at is a more realistic picture of adoption than uniform use.
The authors' own conclusion is unusually direct. They write that sustained clinician engagement, rather than algorithmic accuracy, may be the key barrier to effective clinical AI use in emergency departments, and that these findings inform randomised trial design but do not justify clinical deployment of AI clinical decision support at this stage [s1].
A system rated clinically appropriate on 99 of 100 outputs that did not change length of stay and that two thirds of clinicians stopped using is a clean demonstration that output quality and clinical benefit are separate variables [s1].
Primary care in Kenya
The second trial is larger, randomised, and set where the evidence gap is widest. Rigorous data on language models in real-world, low-resource clinical settings has been close to absent, and most of what exists comes from high-income academic centres.
This was a pragmatic, cluster-randomised trial in 16 primary care facilities in Kenya, with clinical officers randomised to use the electronic medical record with or without language-model assistance [s2]. The primary outcome was an expert-adjudicated composite of treatment failure events within 14 days of enrolment [s2].
Between 22 April and 16 July 2025, 9,691 patients were enrolled, overseen by 103 clinical officers — 52 in the assisted arm and 51 in the control arm [s2].
Treatment failure occurred in 102 of 4,693 patients (2.2%) in the intervention arm and 94 of 4,654 (2.0%) in the control arm, with an adjusted odds ratio of 0.77 (95% CI, 0.55 to 1.08, P = 0.13) [s2]. The primary outcome did not differ significantly between groups [s2].
The point estimate favours the intervention while the confidence interval crosses 1. The authors' interpretation is that language-model assistance was safe but did not reduce treatment failure within 14 days, and that any benefit, if present, is probably modest [s2]. No serious adverse events were judged related to the intervention, and independent review of adverse events did not identify a safety signal [s2].
Why the safety findings matter more than they look
Both trials return the same pair of results: the systems did not help measurably, and they did not hurt.
The second half is not trivial. A recurring concern about deploying language models in clinical settings is that plausible-sounding wrong answers will propagate into decisions, particularly where clinicians are stretched. Nearly ten thousand patients in Kenyan primary care produced no safety signal [s2], and 1,138 emergency department patients produced no detected adverse events with 99 of 100 sampled outputs judged appropriate [s1].
That is meaningful evidence against the strong version of the harm hypothesis, in two very different health systems. It is not evidence of benefit, and the two should not be conflated.
What these trials do to the deployment case
They complicate it in a specific way. The usual objection to clinical AI has been accuracy; the usual response has been better models. Both trials suggest the binding constraint is elsewhere.
In the emergency department the system was accurate and got abandoned [s1]. In Kenyan primary care it was used and did not move a hard clinical outcome [s2]. Neither problem is solved by a better model. One is a workflow and attention problem, the other suggests the baseline — a clinical officer with an electronic record — was already good enough that a 2.0% treatment failure rate left little room.
The 2.0% control-arm failure rate is worth dwelling on [s2]. Interventions are easiest to demonstrate where current performance is poor. A trial that starts from 2.0% needs to be very large to detect improvement, which is part of why the confidence interval is wide.
What to watch
Whether anyone runs the randomised trial the emergency department study says its findings inform, and whether it measures sustained adoption as an endpoint rather than an assumption [s1]. Whether the Kenya result reproduces in settings with higher baseline treatment failure, where there is more room to improve [s2]. And whether procurement decisions cite trials like these or continue to cite benchmark scores.
Sources
- [s1] Prospective evaluation of a large language model clinical decision support system in the emergency department, Nature Medicine, online 2026-08-19.
- [s2] Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial, Nature Medicine, 2026;32(8):3032–3039.
Sources
- Prospective evaluation of a large language model clinical decision support system in the emergency department — Nature Medicine , August 19, 2026
- Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial — Nature Medicine , August 1, 2026
Phone screening after tuberculosis matched home visits overall, and missed recurrences
A trial in India found calling tuberculosis survivors was non-inferior to visiting them, on a combined measure. Split into survivors and their contacts, the picture reverses for the group at highest risk.
A language model restaged 51,242 cancer patients from old radiology reports
Researchers built an AI pipeline to convert decades of narrative reports into one modern TNM staging system, reaching 95% accuracy for tumour classification in an expert-labelled test set.
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
An app matched human coaches in a diabetes trial. Both worked about a third of the time.
The headline finding is noninferiority. The more useful finding is what the shared denominator was — and how wide a gap the trial was designed to tolerate.