ANALYSIS

A hospital chatbot missed its primary endpoint, then projected $146,000 in savings

Singapore General Hospital randomised residents to work with and without an LLM assistant. Documentation time fell by 1.82 minutes, which was not significant. The economic model used the point estimates anyway.

Most published evaluations of large language models in medicine are retrospective: run the model over old notes, score the output, report a benchmark. A trial published this month did something rarer — it put a chatbot into a working preoperative clinic, randomised the days it was available, and timed what happened [s1].

The honest summary is that it missed its primary endpoint and found real effects in subgroups. The more interesting part is what the paper then did with those subgroup numbers.

The trial

PEACH — Perioperative AI CHatbot — is an LLM-based clinical decision support system built for preoperative consultations [s1]. Sixteen eligible resident physicians at Singapore General Hospital were screened and 14 recruited; one declined and one was on emergency medical leave [s2]. Participants were stratified by experience — novice, new to the institution, or experienced — and worked a randomised crossover over two consecutive days, alternating between standard care and access to PEACH [s2].

Across the study period 272 patient encounters were recorded: 135 on days without PEACH and 137 on days with it [s2]. PEACH was actually used in 111 of the 137 intervention-day cases, a utilisation rate of 81.0%; the 26 unused encounters were analysed with the control group, giving 161 control and 111 intervention records [s2]. Outputs arrived in 10 to 15 seconds on average [s2].

The underlying model was Claude 3.5 Sonnet [s2].

The primary result

Total consultation time — from the patient entering the room to the note being finished — did not differ: 40.04 minutes with PEACH against 40.66 without (p=0.787) [s2]. Documentation time, the portion spent writing the note, trended lower with PEACH at 17.53 against 19.35 minutes, a difference of 1.82 minutes (95% CI −4.54 to 0.90) that was not statistically significant (p=0.192) [s2].

That is the headline result, and the paper reports it plainly [s1].

Where it did move

Stratified by case complexity, PEACH was associated with a mean reduction of 5.77 minutes per patient in complexity-2 cases (95% CI −10.14 to −1.41; p=0.010) [s2]. In complexity-1 cases the reduction was 2.51 minutes (p=0.164) and in complexity-3 cases 0.09 minutes (p=0.980) — neither significant [s2].

Stratified by experience, documentation time fell among experienced residents, from 24.6 to 20.0 minutes (difference −4.62, 95% CI −9.02 to −0.23; p=0.040) [s2]. Novices showed a non-significant 1.21-minute reduction and residents new to the institution a non-significant 1.88-minute increase [s2].

The pattern is coherent — a tool that drafts structured summaries helps most where there is a moderate amount of structure to summarise, and helps people who already know what a good note looks like — but it emerged from subgroup analysis of a trial whose overall result was null, and should be treated accordingly.

A confounder worth naming

Usage was not randomised at the encounter level; residents chose when to use the tool on intervention days. The cases where PEACH was used had a significantly higher proportion of high-complexity cases than controls (25.2% versus 13.7%, p=0.007) and more ASA physical status 3 patients (48.6% versus 34.2%, p=0.050) [s2].

So the intervention group was systematically sicker and more complex than the comparison group. That biases against finding a time saving in the overall comparison, and it complicates the complexity-stratified analysis too, since the strata were populated differently by a non-random process.

Usage was also concentrated: among the 14 residents, individual usage rates ranged from 36.4% to 100% of eligible encounters, and two residents accounted for more than 25% of all messages logged [s2].

Quality of the output

Blinded human evaluators compared 28 documents from each arm. They preferred PEACH-assisted documentation in 16 of 28 cases (57.1%) against 10 of 28 (35.7%) for control, which was not statistically significant (p=0.180); inter-rater agreement was substantial (Cohen's κ=0.71) [s2]. PEACH notes were more likely to include an issues list, 20 of 28 (71.4%) against 12 of 28, at p=0.05 [s2].

A note on that last figure: the paper reports 12 of 28 as 43.9%, but 12 divided by 28 is 42.9% [s2]. The raw counts are what matter and they are unambiguous; the derived percentage is a typographical or arithmetic slip.

Minor errors appeared in 3 of 28 control documents (10.7%) and 2 of 28 PEACH documents (7.1%); major errors in 1 control document (3.6%) and none with PEACH [s2]. In a separate review of 30 outputs, all were judged accurate, three (10.0%) contained minor clinical deviations judged not to cause patient harm, and no hallucinations were observed [s2].

Across 168 PEACH interactions in the 111 cases, 71.2% involved a single interaction; outputs were mostly summaries and management plans (82.7%), then question-and-answer responses (11.9%) and referral drafting (5.4%) [s2]. User ratings were high for safety (4.94), explainability (4.81) and ease of understanding (4.72), with usefulness at 4.23 [s2].

The economic model, and its assumptions

The paper projects annual institutional savings of SGD 197,501 (USD 146,297), built from roughly 1,091.4 saved resident hours and 59 saved attending hours per year [s1][s2]. Sensitivity analyses put net annual savings between SGD 48,979 (USD 36,280) in a worst case and SGD 197,499 (USD 146,295) [s1].

Read the inputs. The model assumes 20,000 cases per year, distributed across complexity strata (8,900 / 7,420 / 3,680), and applies per-case time savings of 2.51, 5.77 and 0.09 minutes respectively [s2]. Two of those three figures come from comparisons that were not statistically significant [s2]. Labour costs are SGD 140 per hour for residents and SGD 200 for attendings, with 20% overheads; the model's LLM inference cost is SGD 14.19 per year [s2].

An annual inference bill of about fourteen Singapore dollars against six figures of projected labour savings tells you that the entire result is a function of the assumed minutes, not of the technology's cost. If the true overall documentation saving is the trial's non-significant 1.82 minutes — or zero — the model's output changes accordingly.

The worst-case scenario the authors model assumes 50% adoption, 20% lower time savings, a 50% increase in token price and SGD 30,000 of annual IT maintenance [s2]. It does not model the scenario in which the primary endpoint's null result is the true effect.

What it establishes

That an LLM assistant can be deployed in a high-volume perioperative clinic, used voluntarily in 81% of eligible encounters, produce output that reviewers found accurate and slightly preferred, and generate no observed hallucinations in the sample reviewed [s2]. Those are meaningful operational findings, and the authors describe this as the first prospective real-world evaluation of an LLM-powered decision support tool in perioperative medicine [s2].

What it does not establish is that it saves time overall, in a single institution, over two days, with 14 residents. The authors' own conclusion is appropriately conditional — that PEACH may enhance documentation efficiency and offer economic value [s1].

What to watch

Whether a longer trial with encounter-level randomisation reproduces the complexity-2 finding, and whether the experienced-resident effect is about skill or about familiarity with the local documentation template. Both are testable; neither is settled here.

Sources

Sources

  1. Clinical and economic impact of a large language model in perioperative medicine: a randomized crossover trialnpj Digital Medicine , July 21, 2025
  2. Clinical and economic impact of a large language model in perioperative medicine — open-access full text and tablesEurope PMC (PMC12280117) , July 21, 2025
Related coverage