ANALYSIS

Physicians got 20 points more accurate on rare bone diseases with AI help

42 orthopedic physicians diagnosed 40 rare diseases twice, once alone and once after seeing AI suggestions. Accuracy jumped 20 to 26 points, but the same-day design leaves memory unaccounted for.

Most studies of AI in diagnosis test either how accurate the AI is on its own, or how accurate physicians are on their own — rarely both, in a design that shows what happens when a physician actually sees an AI's suggestion before making the call. A study published this month in the Journal of Medical Internet Research did both, specifically for a category of disease where diagnosis is notoriously hard: rare orthopedic conditions [s1].

Why rare orthopedic diseases are a hard test case

Orthopedic-related rare diseases are difficult to diagnose because of their low prevalence, how differently they can present from patient to patient, and how fragmented the specialist knowledge base is [s1]. That combination makes them a meaningful stress test for whether AI assistance actually helps a working physician, rather than just performing well on a benchmark.

How the study was built

Researchers selected 40 orthopedic-related rare diseases from the Chinese Rare Disease Catalog [s1]. First, four general-purpose large language models each generated a primary diagnosis and five differential diagnoses for every case, with accuracy compared statistically across models [s1]. Then, researchers integrated a representative model into a two-stage workflow tested on 42 practicing orthopedic physicians — 27 classified as intermediate-level and 15 as senior [s1]. Physicians diagnosed all 40 cases independently first, then rediagnosed the same cases after reviewing the AI's non-authoritative suggestions [s1]. Afterward, physicians completed an 8-item questionnaire on how they felt about the workflow [s1].

How the AI models compared to each other

On their own, three of the four models — Claude Sonnet 4.5, ChatGPT-5.0, and Gemini 2.5 Pro — each reached 90% primary-diagnosis accuracy (36 of 40 cases) [s1]. The fourth, DeepSeek-V3.2, scored notably lower at 67.5% (27 of 40), a statistically significant gap (P<.001) [s1].

How much physicians improved with AI help

This is the study's central finding. Before seeing any AI suggestions, intermediate physicians averaged 42.22% diagnostic accuracy and senior physicians averaged 58.67% [s1]. After reviewing AI suggestions on the same cases, those numbers rose to 68.80% for intermediate physicians and 83.33% for senior physicians [s1] — gains of roughly 26 and 25 percentage points, respectively. Both physician seniority and the AI-assistance stage were independently, significantly associated with diagnostic correctness (both P<.001) [s1]. Looking at it by case rather than by individual diagnosis, the number of cases correctly diagnosed by a majority of physicians rose similarly, and the number of cases correctly diagnosed by all three top-performing AI models rose from 16 to 27 out of 40 [s1].

Notably, seniority still mattered even with AI help: senior physicians retained higher accuracy than intermediate physicians throughout, and the gap between the two groups didn't statistically disappear (the seniority-by-stage interaction wasn't significant, P=.10) [s1] — meaning AI assistance boosted both groups, but didn't erase the underlying experience gap.

How physicians felt about it

The post-study questionnaire showed generally positive attitudes toward the AI-assisted workflow, with high internal consistency in responses (Cronbach's alpha of 0.902) and no significant difference in attitudes between the intermediate and senior physician groups [s1].

The caveat the authors themselves emphasize

The study's authors are direct about its central limitation: physicians diagnosed the same 40 cases twice on the same day, first without and then with AI assistance [s1]. That same-day, repeated-case design means some of the accuracy improvement could reflect short-term recall — physicians remembering details from their first pass, or being primed to reconsider a case they'd just seen — rather than the AI assistance itself [s1]. The authors explicitly call their findings exploratory and say they warrant confirmation through more rigorous designs: prospective, randomized, crossover studies with a washout period and independent (not repeated) cases [s1].

What to take from this, honestly

This is one of a relatively small number of AI-in-medicine studies that actually measured physician behavior change rather than just model accuracy — a meaningfully different and more clinically relevant kind of evidence. The size of the reported improvement is large enough to be notable even accounting for the study's own acknowledged limitation. But that limitation is real, and the authors' own call for a properly controlled follow-up study, with a washout period and fresh cases rather than repeated ones, is the right standard to hold this finding to before treating a 25-point accuracy jump as an established effect of AI assistance itself.

Sources

Sources

  1. Diagnostic Performance of Large Language Models for Orthopedic-Related Rare Diseases and Their Impact on Physicians' Diagnostic Accuracy: 2-Stage Comparative Evaluation Study Based on the Chinese Rare Disease CatalogJournal of Medical Internet Research , July 24, 2026
Related coverage