AI chatbots ace medical scenarios alone. Give them to real people, and accuracy collapses.
A 1,298-person randomized study found people using chatbots to work through medical scenarios did no better than people without them — even though the same chatbots, tested alone, got the right answer most of the time.
As chatbots built for fielding medical questions keep multiplying, a randomized study published in Nature Medicine offers a blunt reality check: the gap between how well an AI model performs a medical task in isolation and how well an actual person performs that same task with the AI's help can be enormous [s1].
The study
Researchers tested whether large language models could help members of the public work through ten medical scenarios — identifying an underlying condition and deciding what to do about it — in a controlled trial with 1,298 participants [s1]. Participants were randomly assigned to get help from one of three chatbots (GPT-4o, Llama 3, or Command R+) or to use whatever resource they'd normally turn to, as a control group [s1].
Tested on their own, without a human in the loop, the chatbots performed well: they identified the correct underlying condition in 94.9% of cases on average, and picked the right course of action (what the study calls "disposition" — whether to seek emergency care, see a doctor, or manage something at home) in 56.3% of cases [s1].
Handed to actual participants, that accuracy did not survive contact with real users. People using the same chatbots correctly identified the relevant condition in fewer than 34.5% of cases, and picked the right disposition in fewer than 44.2% — no better, statistically, than the control group working without any AI assistance at all [s1].
Why the gap is so large
The researchers point to the interaction itself as the failure point, not the underlying model [s1]. Standard ways of evaluating medical AI — benchmark exam questions, or simulated conversations with actors playing patients — did not predict the breakdown the study found once real, non-expert users were doing the asking, interpreting the answers, and deciding what to do with them [s1]. How a person describes their symptoms, which follow-up questions they think to ask, and how they read an AI's hedged or partial answer all shape the outcome in ways a benchmark score doesn't capture.
The study's authors frame this as a methodological problem for the field broadly: systems can look reliable under the testing conditions vendors typically publish, and still fail once ordinary users are the ones driving the conversation [s1]. Their recommendation is that developers run systematic human user testing — not just model benchmarks or scripted patient simulations — before deploying these tools for public use [s1].
Why it matters now
The study lands in the middle of a wave of consumer-facing AI health tools reaching the public — chat assistants built into phones, wearables, and search engines that field symptom questions directly from consumers rather than through a clinician. None of that expansion is addressed by this paper directly, and the study does not name or test any specific consumer product beyond the three general-purpose models listed [s1]. What it does establish, with a sample size and randomized design that most AI health claims lack, is that a chatbot's demonstrated competence on a medical question, tested by itself, is not evidence of how well an ordinary person will do using that same chatbot to make a decision about their own health.
What the study doesn't tell us
The ten scenarios were fixed and scripted, which is a limitation the authors themselves would recognize — real symptoms don't arrive as a clean vignette. The three models tested (GPT-4o, Llama 3, Command R+) are specific, dated systems, and newer models may perform differently; the study makes no claim about any model released after its own testing. And "disposition" accuracy in the 40–56% range, even for the models tested alone, means these tools were already imperfect at the one task — telling someone whether they need emergency care — where the cost of being wrong is highest [s1].
Sources
- Reliability of LLMs as medical assistants for the general public: a randomized preregistered study — Nature Medicine, 2026-02-09
Sources
- Reliability of LLMs as medical assistants for the general public: a randomized preregistered study — Nature Medicine , February 9, 2026
A consumer AI health tool undertriaged half the emergencies in a stress test
Researchers ran 960 responses through ChatGPT Health using clinician-written vignettes. Failures clustered at both extremes, and crisis safeguards activated unpredictably.
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
Most AI models that predict who will skip their medicines aren't ready for the clinic
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.
Two FDA-cleared AI tools read prostate biopsies. A review flags what could go wrong.
A review of Paige Prostate Detect and Ibex Prostate Detect finds the tools mainly help less-specialized pathologists, and warns performance shifts when a tool meets a new population.