AI reads chest X-rays for TB about as well as radiologists — and misses about as many
In a prospective African validation, TB-detection AI matched radiologists but neither hit the WHO sensitivity target. In an Indian field programme, AI-led screening lifted case-finding sharply.
Software that reads chest X-rays for signs of tuberculosis performs about as well as radiologists as a triage test — but in the population where it was most rigorously validated, neither the AI nor the radiologists met the World Health Organization's own sensitivity target [s1]. That is the honest state of a technology now being deployed to find TB cases in places with few radiologists, where a real Indian field programme reported a large jump in case-finding after AI-led screening [s2].
The validation that anchors the field
The most rigorous prospective test comes from a study published in NEJM AI in 2024, run in a population with a high burden of TB and HIV across three clinical sites [s1]. It recruited 1,978 adults who had TB symptoms, were close contacts of TB patients, or were newly diagnosed with HIV [s1]. Of the 1,910 patients analysed, 1,827 (96%) had a conclusive TB status, of whom 649 (36%) were HIV-positive and 192 (11%) were TB-positive [s1]. Two cloud-based AI systems were assessed — one to detect TB, one to flag any chest X-ray abnormality — against a microbiological reference standard and against ten radiologists reading the same images blinded [s1].
It matters who built the tools: the study was conducted by the developer, Google, with an independent clinical partner in Zambia, and used an independent laboratory reference standard and a blinded human comparator [s1]. That design is stronger than a vendor accuracy claim, but it is not the same as a fully independent evaluation, and the results should be read with the authorship in view.
The numbers, and the target they missed
At a high-sensitivity operating threshold, the TB AI reached 87% sensitivity and 70% specificity; at a balanced threshold tuned to resemble a radiologist, it reached 78% sensitivity and 82% specificity [s1]. The ten radiologists averaged 76% sensitivity and 82% specificity [s1]. So at the high-sensitivity setting the AI was statistically non-inferior to radiologists on sensitivity (p<0.001) but not on specificity (p=0.99); at the balanced setting it was comparable to radiologists on both [s1].
The finding that should temper any enthusiasm is the WHO benchmark. WHO's targets for a TB triage test are 90% sensitivity and 70% specificity [s1]. The AI cleared the specificity target but not the sensitivity one — and neither did the radiologists [s1]. In a high-HIV population, where TB often looks atypical on an X-ray, both machine and human reading left roughly one in eight to one in seven confirmed cases uncaught at the tuned threshold. A chest X-ray, AI-read or not, is a triage step that decides who gets a confirmatory molecular test, not a diagnosis; missed cases at this stage are missed unless something downstream recovers them. The separate abnormality-detecting AI did clear its easier triage targets, at 97% sensitivity and 79% specificity [s1].
What deployment looks like
A prospective programme published in Open Forum Infectious Diseases in January 2026 shows the field version of this. Working across three public health facilities in tribal districts of Chhattisgarh, India, researchers used the qXR software, made by Qure.ai, to screen for TB [s2]. Of 2,745 chest X-rays screened, 363 patients were identified as presumptive for TB, and among those tested the TB positivity rate was 44.63% (95% CI 39.44–49.91) [s2]. During the AI implementation period, TB case notifications rose by 80.21% against the baseline period, from 96 to 173 cases [s2].
Two cautions belong next to those numbers. First, the study did not compute the software's own sensitivity or specificity against a reference standard; it cites the algorithm's previously reported validation of 0.93 sensitivity and 0.75 specificity for culture-confirmed TB, which is a developer figure from earlier studies, not a finding of this deployment [s2]. Second, several authors are employees of Qure.ai, and the work was grant-funded through a TB philanthropic fund [s2]. The programme is best read as evidence that AI-led screening can raise case-finding where radiologists are scarce — a real and useful result — rather than as an independent accuracy check.
What it means
For a health system choosing whether to put AI at the front of a TB screening pathway, the two studies say complementary things. Where radiologists are unavailable, an AI reader can match the performance a radiologist would give and process X-rays at scale [s1], and in the field that can translate into substantially more cases found and notified [s2]. But the ceiling is the test itself: chest imaging, however it is read, does not reliably reach WHO's sensitivity target in high-HIV populations, so AI raises access and throughput more than it raises the underlying accuracy of X-ray triage [s1]. The cases it flags still need a molecular confirmatory test, and the cases it misses still need another route to diagnosis.
For related coverage, see our reporting on the FDA's testing gap for AI radiology devices, two 2026 radiology-assist studies, phone versus home TB screening after treatment, and the global TB report's funding gap.
Sources
- Prospective Multi-Site Validation of AI to Detect Tuberculosis and Chest X-Ray Abnormalities — NEJM AI, 2024-09-26
- Evaluating the Usefulness of Artificial Intelligence-based Chest X-Ray Screening in Improving Tuberculosis Detection Among the High-Risk Tribal Population of Chhattisgarh, India — Open Forum Infectious Diseases, 2026-01-07
Sources
- Prospective Multi-Site Validation of AI to Detect Tuberculosis and Chest X-Ray Abnormalities — NEJM AI , September 26, 2024
- Evaluating the Usefulness of Artificial Intelligence-based Chest X-Ray Screening in Improving Tuberculosis Detection Among the High-Risk Tribal Population of Chhattisgarh, India: A Prospective Multi-Centre Study — Open Forum Infectious Diseases , January 7, 2026
More on
AI mammography held interval-cancer rates and raised sensitivity in the MASAI trial
The first randomised trial to report interval cancers found AI non-inferior to double reading, with higher sensitivity at equal specificity. Across three European programmes the extra detection is about one per 1000.
Phone screening after tuberculosis matched home visits overall, and missed recurrences
A trial in India found calling tuberculosis survivors was non-inferior to visiting them, on a combined measure. Split into survivors and their contacts, the picture reverses for the group at highest risk.
TB deaths fell again in 2024 — and the money to keep them falling is not there
WHO's Global Tuberculosis Report 2025 counts 10.7 million people ill and over 1.2 million dead last year. Available funding is US$5.9 billion against a US$22 billion annual target.
Autonomous AI can screen for diabetic retinopathy — the catch is the images it can't read
Systems that give a diagnosis with no clinician in the loop post high sensitivity in trials. In real-world use, one in six images was too poor to grade — and the strongest numbers come from company-run studies.