WHAT THE STUDY ACTUALLY SAYS

AI read cervical-screening images with strong accuracy — inside its own datasets

Two 2026 studies posted high accuracy for AI reading Pap-smear cells and colposcopy images in settings short of pathologists. Both were validated internally, which is where such numbers usually shrink.

AI tools for cervical screening take the images a programme already produces — the stained cells of a Pap smear, or a colposcopy picture of the cervix — and classify them as normal or abnormal, standing in for the pathologists and colposcopists that many low- and middle-income countries do not have enough of. Two 2026 studies reported strong accuracy for that task [s1][s2]; both were validated inside their own datasets, which is precisely the evidence that tends to shrink when a tool meets a new clinic and a different population.

The problem these tools target

Cervical cancer remains a leading cause of cancer death among women in sub-Saharan Africa, and a critical shortage of trained pathologists limits how quickly abnormal results can be found and acted on [s1]. Cytology — reading Pap-smear cells under a microscope — has historically driven prevention, but it is labour-intensive and subject to observer variability and limited sensitivity, which is why many programmes now use HPV testing as the primary screen and reserve cytology to triage HPV-positive women [s3]. AI is being pitched as the way to keep image-reading available where specialists are scarce.

The Pap-smear benchmark

The first study is a retrospective analysis that trained five convolutional neural networks — EfficientNetB7, MobileNet, ResNet50, ResNet152 and InceptionNetV3 — to sort cervical cells into six classes, from normal through the graded abnormalities to carcinoma [s1]. It used 11,955 images from the publicly available Center for Recognition and Inspection of Cells (CRIC) database and tested on a held-out set of 961 images [s1].

EfficientNetB7 performed best, with a macro F1 score of 0.9324 (95% CI 0.920–0.945), an accuracy of 0.9775 (95% CI 0.968–0.987) and a sensitivity of 0.9324 [s1]. The carcinoma class was recognised almost perfectly, with recall of 0.978 or higher across all five models [s1]. The models' main confusion was between the two lower grades of abnormality, ASC-US and LSIL, which overlap cytologically [s1]. The important qualifier is what this is: a model tested on a curated public image library, not a clinic. The authors themselves conclude that future work should train on locally representative data and analyse whole smears rather than pre-selected cells [s1].

The real-world colposcopy study

The second study, CITOBOT, is closer to a working programme. It evaluated an AI system on colposcopy images from 650 women screened at a public hospital in Cali, Colombia, between February 2023 and July 2025, with colposcopy-guided biopsy as the reference standard [s2]. Across 2,648 images classified as screen-negative or screen-positive, the system reached an accuracy of 94.3%, sensitivity of 93.4%, specificity of 94.9% and an area under the curve of 0.98, with a positive predictive value of 92.9% and a negative predictive value of 96.2% [s2].

Those are strong numbers against a biopsy standard — but the study is explicit that the performance is internally validated, estimated with five-fold cross-validation and a hold-out subset drawn from the same dataset [s2]. The authors also flag an oddity in their own results: HPV status showed an unexpected inverse association with an AI screen-positive result, which they caution should not be read as biologically protective, since HPV status was never an input to the model [s2]. Reporting that rather than burying it is the mark of a careful analysis.

What the caveat amounts to

A 2026 review of AI in cervical cytology reaches the conclusion both studies point toward: AI-assisted systems may improve efficiency, standardisation and consistency and reduce workload in resource-constrained settings, but the evidence is heterogeneous, and prospective validation, regulatory approval, digital infrastructure and workflow integration remain unresolved [s3]. Its bottom line is that AI should be regarded as an adjunct to human expertise, not a replacement, in cervical screening pathways [s3].

What it means, and what to watch

For programmes chasing the WHO goal of eliminating cervical cancer, an AI that can triage images where no pathologist is available is a real prospect — but the 2026 evidence describes potential, not deployed screening performance. Both headline results come from internal validation, the setting that most flatters an algorithm; the number that matters is how they hold up prospectively, in a new clinic, on locally collected images. That is the same domain-shift problem that recurs across AI-in-medicine, and it is why the honest reading is "promising, not proven."

For the broader screening picture, see our coverage of the shift to HPV as the primary cervical screen, what the HPV vaccine has done to cervical-cancer rates, and the WHO's cervical precancer training push. For a parallel low-resource AI-triage story, see AI chest X-ray screening for tuberculosis.

Sources

Sources

  1. Diagnostic accuracy of an artificial intelligence-driven cytopathological tool in detecting cervical pre-cancerous lesions from Pap smear images — Frontiers in Digital Health , September 3, 2026
  2. CITOBOT AI for real-world cervical cancer screening using colposcopy imaging — Frontiers in Public Health , June 12, 2026
  3. Artificial Intelligence in Cervical Cytology: Opportunities and Limitations in Screening, Triage, and Diagnostic Support — Diagnostics , May 19, 2026

More on

Related coverage