Symptom-checker apps get the diagnosis first about a third of the time
The landmark audit found the right diagnosis listed first in 34% of cases and correct triage in 57%. A newer heart-attack test exposes the dangerous edge: most atypical cases missed, and women.
| Group | Value (%) |
|---|---|
| Emergent care needed | 80 (75 to 86) |
| Non-emergent care | 55 (47 to 63) |
| Self-care reasonable | 33 (26 to 40) |
Symptom-checker apps — the tools that ask what hurts and return a list of possible causes and a "see a doctor" verdict — put the correct diagnosis at the top of the list about a third of the time, and give appropriate triage advice a little over half the time [s1]. Those are the headline numbers from the defining audit, and a newer, harder test shows they degrade in exactly the situation where a wrong answer is most dangerous: an unusual presentation of a life-threatening condition [s2].
The benchmark audit
The reference study fed 45 standardised patient vignettes — evenly split between emergencies, non-emergencies and self-care situations — into 23 publicly available symptom checkers, generating hundreds of evaluations [s1]. For diagnosis, the checkers listed the correct condition first in 34% of evaluations (95% CI 31% to 37%) and somewhere in their top 20 suggestions in 58% (55% to 62%) [s1]. For triage — the more consequential output, since it tells a person whether to go to hospital — they gave appropriate advice in 57% of cases (52% to 61%) [s1].
The average conceals the important structure. Triage accuracy tracked the stakes in the wrong direction for reassurance but the right direction for safety: the checkers were most cautious about genuine emergencies, advising appropriate emergent care in 80% of those cases (75% to 86%), but correct in only 55% of non-emergent cases (47% to 63%) and 33% of self-care cases (26% to 40%) [s1]. In other words the tools err toward telling people to seek care — safer for the individual, but it undercuts the promise of diverting unnecessary visits, and it still misses one emergency in five [s1]. Performance ranged widely between individual checkers, from 33% to 78% appropriate triage [s1].
Where it gets dangerous
A decade later, a 2024 study narrowed the lens to a single condition where errors kill: myocardial infarction [s2]. Researchers entered the anonymised symptoms and biodata of 100 consecutive real heart attack patients into eight commercial symptom checkers [s2]. Overall, a checker named heart attack as its single top diagnosis in 48% of cases (with enormous spread between apps, from 6% to 85%), reached it within the top three in 73%, and within the top five in 79% [s2]. Triage sensitivity — flagging the case as needing urgent treatment — was better, at 83% [s2].
The failure lived in the atypical cases. Where the presentation did not follow the textbook, the top diagnosis was correct in only 24% of cases, and for atypical presentations in women that figure fell to 10% [s2]. Atypical triage sensitivity was significantly worse than for typical cases (53% versus 84%) [s2]. Because women more often present atypically, the tools were significantly less accurate for them — a bias that reproduces one of the oldest failures in cardiology, now encoded in software [s2].
The safety-net question
This is the core problem with self-triage tools: not that they are wrong on average, but that their errors are not evenly distributed, and the reassuring "false negative" — telling someone their emergency is nothing — is the one that harms. A tool that is 80% sensitive to emergencies still sends one in five false comfort [s1]. The design response the literature keeps returning to is safety-netting: erring toward caution, and pairing any low-urgency output with explicit advice on the warning signs that should override it. Most consumer checkers do the first, imperfectly, and the second inconsistently.
The same weakness shows up in the newer generation of large-language-model tools. A consumer AI health assistant, tested on the same triage task, undertriaged half the emergencies in a structured stress test; and where symptom checkers have been deployed as a health system's digital front door, the measured gains were in engagement more than outcomes. Better underlying models may raise the averages, but they do not by themselves fix the distribution of errors, which is what safety depends on.
What it means for a reader
Symptom checkers are reasonable for orienting a non-urgent question and poor as a final arbiter of a serious one. The evidence says to treat a low-urgency verdict as a hypothesis, not a clearance — and to weight it least where it is weakest: unusual symptoms, and, on the heart-attack data, women [s2]. When a checker does say to seek urgent care, the audit data suggest that advice is usually worth taking [s1].
Sources
- [s1] Evaluation of symptom checkers for self diagnosis and triage: audit study — BMJ (2015)
- [s2] Evaluating the diagnostic and triage performance of digital and online symptom checkers for myocardial infarction — PLOS Digital Health (2024)
Sources
- Evaluation of symptom checkers for self diagnosis and triage: audit study — BMJ , July 8, 2015
- Evaluating the diagnostic and triage performance of digital and online symptom checkers for myocardial infarction — PLOS Digital Health , August 5, 2024
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
A consumer AI health tool undertriaged half the emergencies in a stress test
Researchers ran 960 responses through ChatGPT Health using clinician-written vignettes. Failures clustered at both extremes, and crisis safeguards activated unpredictably.
A language model restaged 51,242 cancer patients from old radiology reports
Researchers built an AI pipeline to convert decades of narrative reports into one modern TNM staging system, reaching 95% accuracy for tumour classification in an expert-labelled test set.
Most AI models that predict who will skip their medicines aren't ready for the clinic
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.