ANALYSIS

Symptom-checker apps get the diagnosis first about a third of the time

The landmark audit found the right diagnosis listed first in 34% of cases and correct triage in 57%. A newer heart-attack test exposes the dangerous edge: most atypical cases missed, and women.

Symptom checkers giving appropriate triage advice, by urgency of conditionEmergent care needed: 80%; Non-emergent care: 55%; Self-care reasonable: 33%0%45%90%Emergent care needed80%Non-emergent care55%Self-care reasonable33%
Symptom checkers giving appropriate triage advice, by urgency of condition
GroupValue (%)
Emergent care needed80 (75 to 86)
Non-emergent care55 (47 to 63)
Self-care reasonable33 (26 to 40)
Symptom checkers giving appropriate triage advice, by urgency of condition Audit of 23 symptom checkers across 45 standardised vignettes; 532 triage evaluations. Source: BMJ

Symptom-checker apps — the tools that ask what hurts and return a list of possible causes and a "see a doctor" verdict — put the correct diagnosis at the top of the list about a third of the time, and give appropriate triage advice a little over half the time [s1]. Those are the headline numbers from the defining audit, and a newer, harder test shows they degrade in exactly the situation where a wrong answer is most dangerous: an unusual presentation of a life-threatening condition [s2].

The benchmark audit

The reference study fed 45 standardised patient vignettes — evenly split between emergencies, non-emergencies and self-care situations — into 23 publicly available symptom checkers, generating hundreds of evaluations [s1]. For diagnosis, the checkers listed the correct condition first in 34% of evaluations (95% CI 31% to 37%) and somewhere in their top 20 suggestions in 58% (55% to 62%) [s1]. For triage — the more consequential output, since it tells a person whether to go to hospital — they gave appropriate advice in 57% of cases (52% to 61%) [s1].

The average conceals the important structure. Triage accuracy tracked the stakes in the wrong direction for reassurance but the right direction for safety: the checkers were most cautious about genuine emergencies, advising appropriate emergent care in 80% of those cases (75% to 86%), but correct in only 55% of non-emergent cases (47% to 63%) and 33% of self-care cases (26% to 40%) [s1]. In other words the tools err toward telling people to seek care — safer for the individual, but it undercuts the promise of diverting unnecessary visits, and it still misses one emergency in five [s1]. Performance ranged widely between individual checkers, from 33% to 78% appropriate triage [s1].

Where it gets dangerous

A decade later, a 2024 study narrowed the lens to a single condition where errors kill: myocardial infarction [s2]. Researchers entered the anonymised symptoms and biodata of 100 consecutive real heart attack patients into eight commercial symptom checkers [s2]. Overall, a checker named heart attack as its single top diagnosis in 48% of cases (with enormous spread between apps, from 6% to 85%), reached it within the top three in 73%, and within the top five in 79% [s2]. Triage sensitivity — flagging the case as needing urgent treatment — was better, at 83% [s2].

The failure lived in the atypical cases. Where the presentation did not follow the textbook, the top diagnosis was correct in only 24% of cases, and for atypical presentations in women that figure fell to 10% [s2]. Atypical triage sensitivity was significantly worse than for typical cases (53% versus 84%) [s2]. Because women more often present atypically, the tools were significantly less accurate for them — a bias that reproduces one of the oldest failures in cardiology, now encoded in software [s2].

The safety-net question

This is the core problem with self-triage tools: not that they are wrong on average, but that their errors are not evenly distributed, and the reassuring "false negative" — telling someone their emergency is nothing — is the one that harms. A tool that is 80% sensitive to emergencies still sends one in five false comfort [s1]. The design response the literature keeps returning to is safety-netting: erring toward caution, and pairing any low-urgency output with explicit advice on the warning signs that should override it. Most consumer checkers do the first, imperfectly, and the second inconsistently.

The same weakness shows up in the newer generation of large-language-model tools. A consumer AI health assistant, tested on the same triage task, undertriaged half the emergencies in a structured stress test; and where symptom checkers have been deployed as a health system's digital front door, the measured gains were in engagement more than outcomes. Better underlying models may raise the averages, but they do not by themselves fix the distribution of errors, which is what safety depends on.

What it means for a reader

Symptom checkers are reasonable for orienting a non-urgent question and poor as a final arbiter of a serious one. The evidence says to treat a low-urgency verdict as a hypothesis, not a clearance — and to weight it least where it is weakest: unusual symptoms, and, on the heart-attack data, women [s2]. When a checker does say to seek urgent care, the audit data suggest that advice is usually worth taking [s1].

Sources

  • [s1] Evaluation of symptom checkers for self diagnosis and triage: audit study — BMJ (2015)
  • [s2] Evaluating the diagnostic and triage performance of digital and online symptom checkers for myocardial infarction — PLOS Digital Health (2024)

Sources

  1. Evaluation of symptom checkers for self diagnosis and triage: audit studyBMJ , July 8, 2015
  2. Evaluating the diagnostic and triage performance of digital and online symptom checkers for myocardial infarctionPLOS Digital Health , August 5, 2024
Related coverage