ANALYSIS

AI scribes are already in the exam room. Nobody has shown they are safe

A February theme issue documents faster notes, happier clinicians and enterprise rollouts to thousands. It also contains an editorial asking the question none of the studies answer.

An editorial published in JMIR Medical Informatics on 6 February states the position more bluntly than most vendors or health systems have: AI scribes have achieved a pace of clinical adoption rarely seen for digital technologies in health care, and that adoption has occurred well ahead of robust evidence of their safety and efficacy [s1].

The editorial's explanation for the speed is not cynical. The technology works well enough, it addresses a genuine pain point for clinicians, and it has largely sidestepped regulatory requirements [s1]. Documentation times decrease, clinicians report feeling less burdened, and the notes produced are often of reasonable quality [s1]. What remains unanswered, the editorial argues, is the outstanding and urgent question: are AI scribes safe? What are the clinical outcomes achievable when scribes are used, compared with other forms of note taking [s1]?

The studies published alongside it show how far implementation has run.

The deployment data

An academic health system made ambient scribing available to over 2,400 ambulatory and emergency department clinicians simultaneously on 15 January 2025 [s3]. By 31 March 2025, 20.1% of visit notes incorporated ambient scribing and 1,223 clinicians had used it [s3]. Among 209 survey respondents — 22.1% of the 947 surveyed — 90.9% said they would be disappointed if they lost access, and 84.7% reported a positive training experience [s3]. The authors conclude that simultaneous enterprise-wide deployment was feasible and that support needs were manageable [s3].

At UCI Health, where a quality improvement pilot has run since December 2023, electronic health record usage data from 167 physicians showed significant reductions in note-writing time despite an increase in note length [s4]. Matched pre- and post-implementation surveys (n = 65) found statistically significant reductions in reported cognitive demand (P = .031) and documentation effort (P = .014), alongside perceptions of improved clinical efficiency, patient-centred care and EHR usability [s4].

The emergency department picture is more granular, and more revealing about who actually uses these tools. In a retrospective observational study of adult ED encounters at a tertiary academic medical centre, published in Annals of Emergency Medicine on 10 February, attending physicians could optionally use an ambient scribe [s2]. Among 8,740 eligible encounters, 976 — 11.2% — used it [s2]. Thirty-five of 92 attendings (38%) used the tool at all, and a small group of high-frequency users accounted for most ambient encounters [s2].

Where it was used matters. Ambient use clustered in telemedicine and vertical-care zones, in lower-acuity patients, and in encounters not requiring interpreters [s2]. When used, median on-shift documentation time was 2 minutes 45 seconds for ambient encounters versus 3 minutes 50 seconds for standard ones — a difference of 1 minute 5 seconds, or 28% [s2]. Median total EHR time was 8 minutes 39 seconds versus 10 minutes 21 seconds, a 16% reduction, and ambient notes were shorter overall [s2].

That is a real efficiency gain, measured in audit logs rather than self-report. It is also a gain observed in the encounters clinicians selected as suitable — the easier ones — which is not the same as a gain available across the case mix.

What the notes lose

A qualitative study at a large academic medical centre interviewed clinicians from an ambient scribe pilot (n = 8) and the initial enterprise rollout (n = 16), in sessions of 26 to 60 minutes [s5]. Clinicians described feeling more present with patients and greater satisfaction during visits [s5]. They also described overlong or underspecified note sections, unfamiliar formatting, and a perceived loss of "voice" [s5].

The analytic point is the one efficiency metrics miss. Participants discussed using documentation to personalise practice, demonstrate expertise, manage impressions with colleagues and supervisors, and communicate sensitive findings — activities not fully captured by efficiency metrics [s5]. In inpatient and procedure-heavy contexts, where documentation was already highly standardised, benefit was limited [s5]. Early implementation, the authors conclude, introduced new work to reconcile AI-drafted text with local documentation conventions and audience-specific communication [s5].

The safety concern behind the editorial has an empirical anchor. A systematic review and meta-analysis of human–LLM collaboration in clinical medicine, published in npj Digital Medicine on 28 January, found that while documentation quality improved across the studies it pooled, factual error rates remained high — roughly 26% to 36% — which the authors say undermines the quality gains [s7].

Patients have not been asked much

The largest study in the group surveyed patients rather than clinicians: 12,153 adults in Canada, surveyed between 6 February and 10 March 2025, of whom 52.4% were female, 23.1% aged 65 or over, and 41.2% living with chronic conditions [s6].

Awareness of AI scribe use was low, at 28.3% [s6]. Attitudes were mixed: 39.3% reported some or very high comfort, 57.4% trusted documentation with human oversight, and 49.5% anticipated positive effects on patient–provider interactions [s6]. Yet 61.8% were reluctant about future AI scribe use [s6]. Men had higher odds of favourable comfort (aOR 1.13, 95% CI 1.05–1.22, P = .001) and trust (aOR 1.21, 95% CI 1.10–1.32) [s6].

The authors call this a paradox — conditional trust and comfort alongside reluctance to adopt — and identify privacy concerns and low awareness as the key barriers, arguing that targeted work on digital literacy, privacy safeguards and clinician–patient communication is needed before widespread adoption [s6].

The sequencing there is worth noticing. The survey describes what should happen before widespread adoption. The deployment studies describe adoption that has already happened.

Limits

These are single-site or single-system studies, mostly observational, with self-selected users. The ED analysis is retrospective and compares encounters clinicians chose to document differently, not randomly assigned ones [s2]. The survey work is cross-sectional. None of the studies here measures diagnostic accuracy, missed findings, or downstream patient outcomes — which is precisely the editorial's complaint [s1].

What to watch

Whether any study reports clinical outcomes rather than documentation time, which the JMIR editorial names as the missing evidence [s1]. Whether adoption broadens beyond low-acuity, non-interpreted encounters, or stays concentrated where it is easiest [s2]. And whether patient awareness — 28.3% in the largest survey to date — moves at anything like the speed of clinician uptake [s6].

Sources

  1. [s1] AI Scribes: Are We Measuring What Matters? JMIR Medical Informatics, published online 6 February 2026. https://doi.org/10.2196/89337
  2. [s2] Ambient Artificial Intelligence Scribe Adoption and Documentation Time in the Emergency Department. Annals of Emergency Medicine, published online 10 February 2026. https://doi.org/10.1016/j.annemergmed.2025.12.017
  3. [s3] Enterprise-wide simultaneous deployment of ambient scribe technology: lessons learned from an academic health system. Journal of the American Medical Informatics Association, published online 31 October 2025; February 2026 issue. https://doi.org/10.1093/jamia/ocaf186
  4. [s4] Evaluating ambient artificial intelligence documentation: effects on work efficiency, documentation burden, and patient-centered care. Journal of the American Medical Informatics Association, published online 16 October 2025; February 2026 issue. https://doi.org/10.1093/jamia/ocaf180
  5. [s5] Listening to the note: clinician perspectives on ambient artificial intelligence scribes in medical documentation. Journal of the American Medical Informatics Association, published online 3 December 2025; February 2026 issue. https://doi.org/10.1093/jamia/ocaf214
  6. [s6] Patient attitudes toward ambient artificial intelligence scribes in clinical care: insights from a cross-sectional study. Journal of the American Medical Informatics Association, published online 5 December 2025; February 2026 issue. https://doi.org/10.1093/jamia/ocaf218
  7. [s7] Human–large language model collaboration in clinical medicine: a systematic review and meta-analysis. npj Digital Medicine, published online 28 January 2026. https://doi.org/10.1038/s41746-026-02382-2

Sources

  1. AI Scribes: Are We Measuring What Matters?JMIR Medical Informatics , February 6, 2026
  2. Ambient Artificial Intelligence Scribe Adoption and Documentation Time in the Emergency DepartmentAnnals of Emergency Medicine , February 10, 2026
  3. Enterprise-wide simultaneous deployment of ambient scribe technology: lessons learned from an academic health systemJournal of the American Medical Informatics Association , October 31, 2025
  4. Evaluating ambient artificial intelligence documentation: effects on work efficiency, documentation burden, and patient-centered careJournal of the American Medical Informatics Association , October 16, 2025
  5. Listening to the note: clinician perspectives on ambient artificial intelligence scribes in medical documentationJournal of the American Medical Informatics Association , December 3, 2025
  6. Patient attitudes toward ambient artificial intelligence scribes in clinical care: insights from a cross-sectional studyJournal of the American Medical Informatics Association , December 5, 2025
  7. Human–large language model collaboration in clinical medicine: a systematic review and meta-analysisnpj Digital Medicine , January 28, 2026
Related coverage