ANALYSIS

Two August studies point language models at clinical text, with mixed verdicts

An EHR-embedded model wrote hospital summaries residents edited less — but with more confabulations. A separate pipeline audited 21,041 trial reports against CONSORT at 91.7% agreement with experts.

Share of the hospital course text edited by residentsLLM-written: 31.5%; Physician-written: 44.8%0%25%50%LLM-written31.5%Physician-written44.8%
Share of the hospital course text edited by residents
GroupValue (%)
LLM-written31.5
Physician-written44.8
Share of the hospital course text edited by residents 100 admissions, each with a blinded pair of drafts edited for three minutes apiece. Source: JAMA Network Open

Two studies published in JAMA Network Open this month apply large language models to clinical text in very different settings. One asks whether a model can draft the narrative section of a discharge summary well enough that a doctor edits it less than a colleague's draft. The other uses a model as an auditor, checking more than twenty thousand published trials against a reporting standard. Neither produces an unambiguous verdict, and the ways they fall short are instructive.

Drafting the hospital course

The hospital course is the narrative section of a discharge summary describing what happened during an admission. It is time-consuming to write and consequential to get wrong, because the next clinician to see the patient often reads it instead of the chart.

Researchers at New York University Langone Health ran a quality improvement study using a convenience sample of 10 internal medicine resident editors, 8 hospitalist evaluators, and randomly selected general medicine admissions from December 2023 lasting four to eight days [s1]. Residents and hospitalists reviewed assigned records for 10 minutes; residents, blinded to which was which, then edited each pair of hospital courses — one written by a physician, one by an electronic health record-embedded LLM — for three minutes each [s1]. Attending hospitalists then rated the edited pairs head to head [s1].

The evaluation standard was a framework the authors call the 4Cs: complete, concise, cohesive and confabulation-free [s1].

Across 100 admissions, residents edited a smaller percentage of the LLM-written summaries than the physician-written ones — a mean of 31.5% (SD 16.6%) versus 44.8% (SD 20.0%), P < .001 [s1]. The LLM drafts also required less semantic change, meaning less alteration of the original meaning: 2.4% (SD 1.6%) versus 4.9% (SD 3.5%), P < .001 [s1].

On the attending physicians' comparative ratings, the LLM summaries were judged more complete (mean difference 3.00 on a 10-point bidirectional scale, SD 5.28; P < .001), similarly concise (−1.02, SD 6.08; P = .20) and similarly cohesive (0.70, SD 6.14; P = .60) [s1].

And more confabulated: −0.98 (SD 3.53; P = .002) [s1]. The composite scores were similar overall (mean difference 1.70 on a 40-point bidirectional scale, SD 14.24; P = .46) [s1].

What "more confabulations" costs

The completeness advantage and the confabulation disadvantage are probably the same phenomenon viewed from two angles. A model that generates fluent, comprehensive narrative from a chart will fill gaps, and some of what it fills in will not be in the record.

Fewer edits therefore does not straightforwardly mean less work. Catching a fabricated detail in an otherwise excellent paragraph is harder than expanding a terse one, and the study's three-minute editing window is an artificial constraint the authors themselves acknowledge as a potential influence on results [s1].

Their conclusion is appropriately hedged: the study supports the feasibility of a physician-LLM partnership for writing hospital courses and provides a basis for monitoring LLM-generated summaries in clinical practice [s1]. Feasibility and monitoring, not adoption.

The study is also a single-centre quality improvement project with a convenience sample of 10 residents and 8 hospitalists on 100 admissions from one month [s1]. That is a pilot.

Auditing the literature

The second study inverts the relationship: instead of the model producing text for humans to check, humans check whether the model can reliably assess text.

The task was CONSORT compliance. CONSORT is the reporting standard for randomised trials, and the study's premise is that manual audits cannot keep pace with publication volume [s2]. Researchers built a zero-shot pipeline — no task-specific training examples — using GPT-4o-mini to judge whether each of 21 CONSORT items was met [s2].

Validation came first. On a 50-article CONSORT-Text Classification Model benchmark and against expert review of 70 randomly sampled trials, the model's outputs matched experts 91.7% of the time (2,026 of 2,210 decisions), with a macro F1 score of 0.86 (95% CI, 0.84 to 0.87) on the benchmark [s2].

It was then applied at scale. Of 53,137 screened PDFs, 21,041 randomised trials were included — median publication year 2014 (IQR 2003-2020), across 30 disciplines, with a registry-linked subset of 1,790 trials whose median planned enrolment was 210 participants (IQR 95-440) [s2].

What the audit found

Mean CONSORT compliance rose from 27.3% (95% CI, 27.0% to 27.6%) in 1966-1990 to 57.0% (95% CI, 56.8% to 57.2%) in 2010-2024 [s2].

Improvement, then — to just over half. And the items that went unreported are the ones that matter most for judging whether a trial's result can be believed. Allocation concealment mechanism was reported in 16.1% of trials (95% CI, 15.6% to 16.6%) [s2]. Discussion of external validity appeared in 1.6% (95% CI, 1.5% to 1.8%) [s2].

Compliance varied widely by field, from 35.2% in pharmacology (95% CI, 34.8% to 35.6%) to 63.4% in urology (95% CI, 62.1% to 64.7%) [s2]. Associations with trial characteristics — funding source, phase, FDA regulation, oversight features — were negligible, all with Cramér's V below 0.10 [s2].

The last finding is the one that resists easy explanation. Whether a trial was industry-funded, regulated, or subject to formal oversight barely predicted how well it was reported.

What both studies share

Neither claims the model is correct; both measure agreement with humans. The summarisation study measures how much a resident edits and how an attending rates the result [s1]. The audit study measures agreement with expert judgement on a validation set, then extrapolates [s2]. A 91.7% agreement rate applied to 21,041 papers implies a large absolute number of disagreements, and the study does not characterise which items the model gets wrong most often.

Both are also single-model, single-configuration results. Neither tells you what happens when the underlying model is updated — a limitation the CHART reporting guideline for chatbot health advice studies, published on 1 August, was specifically designed to make visible: it asks studies to report model identifiers, model details, prompt engineering and query strategy, among 12 items and 39 subitems [s3].

What they do establish is a division of labour worth taking seriously. Using a language model to draft text a clinician is accountable for carries a confabulation risk that is measurable and non-zero [s1]. Using one to screen a literature nobody currently has capacity to screen at all is a different proposition, where the alternative is not careful human review but no review [s2].

Sources

Sources

  1. Evaluating Hospital Course Summarization by an Electronic Health Record-Based Large Language ModelJAMA Network Open , August 13, 2025
  2. Large Language Model Analysis of Reporting Quality of Randomized Clinical Trial Articles: A Systematic ReviewJAMA Network Open , August 28, 2025
  3. Reporting Guideline for Chatbot Health Advice Studies: The CHART StatementJAMA Network Open , August 1, 2025
Related coverage