The first randomised trials of AI scribes arrived. They do not fully agree.
Two November trials found ambient AI cut documentation time and work exhaustion. A parallel emergency department comparison found physicians spent more time in the note, not less.
Ambient AI scribes — software that listens to a clinical encounter and drafts the visit note — have been deployed across large health systems for two years on the strength of before-and-after evaluations and satisfaction surveys. This month the randomised evidence finally arrived, in the form of two trials published together in NEJM AI, alongside a smaller emergency department comparison in Annals of Emergency Medicine [s1][s2][s3]. The results are more mixed than the deployment pace implies.
Trial one: a stepped wedge in Wisconsin
The first trial was a 24-week, stepped-wedge, individually randomised pragmatic study across ambulatory clinics in two states, enrolling 66 health care practitioners assigned to three six-week sequences of ambient AI [s1]. The co-primary outcomes were professional fulfilment and work exhaustion/interpersonal disengagement, both from the Stanford Professional Fulfillment Index [s1].
Across the trial, 71,487 notes were authored, of which 27,092 — 38% — were generated using ambient AI [s1]. Work exhaustion and interpersonal disengagement fell significantly, by 0.44 points on a five-point Likert scale (95% CI −0.62 to −0.25; P<0.001) [s1]. Professional fulfilment rose by 0.14 points (95% CI 0.004 to 0.28; P=0.04), which the authors characterise as a non-significant increase against their pre-specified threshold [s1].
Time spent on notes fell by 0.36 hours per day (95% CI −0.55 to −0.17) [s1]. A reduction in "work outside work" of 0.50 hours per day (95% CI −0.90 to −0.09) was sensitive to outliers and no longer significant once the top 3% of daily observations were removed [s1]. Documentation quality, scored with the PDSQI-9 instrument, ranged from 3.97 to 4.99 across domains on a five-point scale, and diagnostic billing codes improved (P<0.001) [s1].
Trial two: two products against usual care
The second trial, run at UCLA, randomised 238 outpatient physicians across 14 specialties, 1:1:1, to one of two commercial scribe applications — Microsoft Dragon Ambient eXperience Copilot or Nabla — or usual care, between 4 November 2024 and 3 January 2025 [s2]. Randomisation was covariate-constrained on baseline time-in-note, burnout score and clinic days per week [s2]. The primary outcome was change from baseline in log writing time-in-note [s2].
The two products did not perform alike on the primary endpoint. Nabla users showed a 9.5% decrease in time-in-note versus control (95% CI −17.2% to −1.8%; P=0.02); DAX users showed no significant change (−1.7%; 95% CI −9.4% to +5.9%; P=0.66) [s2]. Uptake was similar — DAX was used in 33.5% of 24,696 visits and Nabla in 29.5% of 23,653 [s2].
On the secondary well-being measures both arms moved in the same direction as the Wisconsin trial: Mini-Z scores rose (DAX +2.83, 95% CI +1.28 to +4.37; Nabla +2.69, 95% CI +1.14 to +4.23 on a 10–50 scale) and physician task load fell (DAX −39.9, 95% CI −71.9 to −7.9 on a 0–400 scale) [s2]. The authors state these secondary findings need confirmation in larger, multicentre trials [s2].
One grade 1 adverse event was reported, and clinicians rated clinically significant inaccuracies as occurring "occasionally" on five-point scales (DAX 2.7; Nabla 2.8) [s2].
The dissenting result
The Annals of Emergency Medicine study is not a randomised trial — it is a quality improvement pilot with five early adopters, comparing ambient AI against human scribes rather than against usual care, from December 2024 to January 2025 [s3]. It covered 710 visits: 284 with human scribes and 426 with AI-assisted charting [s3].
Its findings run the other way. Physicians spent more time in the electronic health record notes section per patient with AI scribes than with human ones — 4.3 versus 1.8 minutes for adults (adjusted risk ratio 2.38, 95% CI 1.85 to 3.05) and 3.5 versus 1.6 minutes for paediatric patients (aRR 2.21, 95% CI 1.94 to 2.51) [s3]. Physicians contributed substantially more of each note themselves with AI: 60.1% of characters versus 30.8% for adults, and 62.3% versus 27.1% for paediatric patients [s3]. Note quality, scored blind on the PDQI-9, was similar for adults but lower for paediatric notes with AI (41.36 versus 42.25, aRR −1.89, 95% CI −3.58 to −0.20) [s3].
Reconciling them
The comparators differ, and that probably explains most of the divergence. The two trials compare ambient AI against a physician typing their own note; the emergency department study compares it against a human scribe already doing the typing. Against no help, AI helps. Against a person, the physician takes back editing work the human scribe had absorbed.
The setting differs too. Emergency medicine involves undifferentiated presentations, frequent interruptions and paediatric encounters where the history often comes from a caregiver — conditions less favourable to a system transcribing a single conversation. The paediatric quality gap is the study's most specific signal [s3].
What none of them settles
All three are short. The longest ran 24 weeks [s1]; the UCLA trial ran two months [s2]; the emergency pilot two months with five physicians [s3]. Burnout measures are self-reported and unblinded — participants know whether they are using a scribe, and that alone can move a satisfaction scale. None of the studies measured patient outcomes, diagnostic accuracy, or downstream consequences of note content. The Wisconsin trial's finding that billing codes improved is reported as a coding outcome, not evaluated as a cost effect [s1].
The most useful thing in the package may be the UCLA result that two products marketed for the same job produced different effects on the primary endpoint [s2]. Ambient AI is not one intervention, and evidence for one vendor's system is not evidence for another's.
Sources
- A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being — NEJM AI, 2025-11-26
- Ambient AI Scribes in Clinical Practice: A Randomized Trial — NEJM AI, 2025-11-26
- Ambient Artificial Intelligence Versus Human Scribes in the Emergency Department — Annals of Emergency Medicine, 2025-11-18
Sources
- A Pragmatic Randomized Controlled Trial of Ambient Artificial Intelligence to Improve Health Practitioner Well-Being — NEJM AI , November 26, 2025
- Ambient AI Scribes in Clinical Practice: A Randomized Trial — NEJM AI , November 26, 2025
- Ambient Artificial Intelligence Versus Human Scribes in the Emergency Department — Annals of Emergency Medicine , November 18, 2025
A computer that grades each colonoscopy raised how often endoscopists found adenomas
A Danish stepped-wedge trial gave endoscopists automated feedback on their technique after every procedure. Adenoma detection rose from 43.4% to 48.6% — a different tool from real-time polyp AI.
More record sharing cut readmissions in community hospitals and raised them inside the VA
A study of 2.4 million veterans a month finds health information exchange volume moving outcomes in opposite directions depending on which side of the exchange you are on.
507 digital therapeutics are approved in four countries. Effect sizes aren't comparable
A new structured dataset finally makes the approvals countable. Two trials published weeks apart show why counting them says little: 0.86 fewer migraine days in one, a 20-point symptom shift in the other.
AI trial-screening tools mostly ignore race. One label skewed results: homelessness.
A study running 5.3 million evaluations through nine large language models found eligibility judgments were largely stable across identity labels — except when a patient vignette mentioned homelessness.