ANALYSIS

AI chatbots ace medical scenarios alone. Give them to real people, and accuracy collapses.

A 1,298-person randomized study found people using chatbots to work through medical scenarios did no better than people without them — even though the same chatbots, tested alone, got the right answer most of the time.

As chatbots built for fielding medical questions keep multiplying, a randomized study published in Nature Medicine offers a blunt reality check: the gap between how well an AI model performs a medical task in isolation and how well an actual person performs that same task with the AI's help can be enormous [s1].

The study

Researchers tested whether large language models could help members of the public work through ten medical scenarios — identifying an underlying condition and deciding what to do about it — in a controlled trial with 1,298 participants [s1]. Participants were randomly assigned to get help from one of three chatbots (GPT-4o, Llama 3, or Command R+) or to use whatever resource they'd normally turn to, as a control group [s1].

Tested on their own, without a human in the loop, the chatbots performed well: they identified the correct underlying condition in 94.9% of cases on average, and picked the right course of action (what the study calls "disposition" — whether to seek emergency care, see a doctor, or manage something at home) in 56.3% of cases [s1].

Handed to actual participants, that accuracy did not survive contact with real users. People using the same chatbots correctly identified the relevant condition in fewer than 34.5% of cases, and picked the right disposition in fewer than 44.2% — no better, statistically, than the control group working without any AI assistance at all [s1].

Why the gap is so large

The researchers point to the interaction itself as the failure point, not the underlying model [s1]. Standard ways of evaluating medical AI — benchmark exam questions, or simulated conversations with actors playing patients — did not predict the breakdown the study found once real, non-expert users were doing the asking, interpreting the answers, and deciding what to do with them [s1]. How a person describes their symptoms, which follow-up questions they think to ask, and how they read an AI's hedged or partial answer all shape the outcome in ways a benchmark score doesn't capture.

The study's authors frame this as a methodological problem for the field broadly: systems can look reliable under the testing conditions vendors typically publish, and still fail once ordinary users are the ones driving the conversation [s1]. Their recommendation is that developers run systematic human user testing — not just model benchmarks or scripted patient simulations — before deploying these tools for public use [s1].

Why it matters now

The study lands in the middle of a wave of consumer-facing AI health tools reaching the public — chat assistants built into phones, wearables, and search engines that field symptom questions directly from consumers rather than through a clinician. None of that expansion is addressed by this paper directly, and the study does not name or test any specific consumer product beyond the three general-purpose models listed [s1]. What it does establish, with a sample size and randomized design that most AI health claims lack, is that a chatbot's demonstrated competence on a medical question, tested by itself, is not evidence of how well an ordinary person will do using that same chatbot to make a decision about their own health.

What the study doesn't tell us

The ten scenarios were fixed and scripted, which is a limitation the authors themselves would recognize — real symptoms don't arrive as a clean vignette. The three models tested (GPT-4o, Llama 3, Command R+) are specific, dated systems, and newer models may perform differently; the study makes no claim about any model released after its own testing. And "disposition" accuracy in the 40–56% range, even for the models tested alone, means these tools were already imperfect at the one task — telling someone whether they need emergency care — where the cost of being wrong is highest [s1].

Sources

Sources

  1. Reliability of LLMs as medical assistants for the general public: a randomized preregistered studyNature Medicine , February 9, 2026
Related coverage