WHAT THE STUDY ACTUALLY SAYS

AI can flag pancreatic cancer on ordinary CT scans, but only in retrospective tests

A deep-learning model reads the pancreas on non-contrast CT taken for other reasons. It reached very high accuracy in large tests — none of them a prospective screening trial.

Pooled accuracy of AI and radiomics models for detecting pancreatic cancerSensitivity: 0.88; Specificity: 0.9300.51Sensitivity0.88Specificity0.93
Pooled accuracy of AI and radiomics models for detecting pancreatic cancer
GroupValue (value)
Sensitivity0.88 (0.84 to 0.91)
Specificity0.93 (0.87 to 0.96)
Pooled accuracy of AI and radiomics models for detecting pancreatic cancer Meta-analysis of 15 studies, 14,688 patients. Sensitivity and specificity are different measures; both carry wide confidence intervals. Source: Journal of Gastrointestinal Surgery

Deep-learning software can now flag pancreatic tumours on a routine non-contrast CT scan — the kind of image taken for an unrelated complaint, without the intravenous dye radiologists normally need to see the pancreas well [s1]. In the largest test to date, one such model matched or beat radiologists and reached near-perfect accuracy on stored scans; what it has not done is prove, in a prospective trial, that using it earlier finds cancers that would otherwise be missed and lengthens lives [s1][s2].

That distinction is the whole story, because pancreatic cancer is the disease where it would matter most.

Why the pancreas is the hard case

Pancreatic ductal adenocarcinoma is the most lethal common solid malignancy, and it is usually found late, at a stage where surgery is no longer possible [s1]. Catching it early or incidentally is associated with longer survival, but screening the general population with any single test has been considered unfeasible: the disease is rare, so even an accurate test throws off large numbers of false positives that lead to anxiety, further scans and unnecessary procedures [s1]. The pancreas is also poorly seen on non-contrast CT, which is why identifying cancer on those scans was long thought to be effectively impossible [s1].

The appeal of an AI reader is that it does not require a new scan. Millions of non-contrast CTs are performed every year for other reasons; a model that can read the pancreas on images already taken offers, in principle, screening at no extra dose, appointment or cost. It is the same logic behind reading coronary calcium off lung-cancer screening scans — extract a second, unrelated finding from an image that already exists. The question is whether the automated number can be trusted.

What PANDA reported

The model that set the benchmark, called PANDA, was published in Nature Medicine in 2023 [s1]. It was trained on scans from 3,208 patients at a single centre [s1]. In a multicentre validation of 6,239 patients across 10 centres, it achieved an area under the curve of 0.986 to 0.996 for detecting a pancreatic lesion [s1]. Against radiologists reading the same scans, it exceeded mean radiologist sensitivity by 34.1 percentage points and specificity by 6.3 points for identifying pancreatic cancer specifically [s1].

The headline figure came from what the authors called a real-world, multi-scenario validation of 20,530 consecutive patients: sensitivity of 92.9% and specificity of 99.9% for detecting a lesion [s1]. The paper also reported that PANDA, using non-contrast CT, was non-inferior to the original radiology reports — which had been written from contrast-enhanced scans — in telling the common types of pancreatic lesion apart [s1].

Those are exceptional numbers, and the 99.9% specificity is the one to sit with. At that level, false positives are rare enough that population screening starts to look conceivable, which is precisely the claim a prospective trial would need to confirm rather than assume.

What the wider literature shows

PANDA is one model. A 2026 systematic review and meta-analysis in the Journal of Gastrointestinal Surgery pooled 15 studies of radiomics and machine-learning models for detecting pancreatic cancer in imaging, covering 14,688 patients, of whom 6,153 (41.8%) had the cancer [s2]. Across those studies the pooled sensitivity was 0.88 (95% CI 0.84–0.91) and specificity 0.93 (95% CI 0.87–0.96) [s2]. The positive likelihood ratio was 12.1 and the negative likelihood ratio 0.12 [s2].

Two things stand out. The pooled specificity of 93% is well short of PANDA's 99.9%, a reminder that a single flagship result is not the field's average — and at the low prevalence of pancreatic cancer, the gap between 93% and 99.9% is the difference between a workable screening tool and one that buries clinicians in false alarms. And the heterogeneity between studies was very high (I² of 87.8% for sensitivity and 95.0% for specificity), meaning the pooled figures average over models and populations that behaved quite differently [s2]. The reviewers' own conclusion is that these results are promising but that prospective studies are needed to establish real-world efficacy [s2].

What is missing

Every result above comes from stored images with the answer already known. None is a prospective screening study in which people are scanned, flagged by the AI, and followed to see whether earlier detection changed what happened to them. That is the evidence that would move a tool from impressive to clinically justified, and it does not yet exist for pancreatic AI on CT.

The gap matters more here than in most imaging tasks. A retrospective dataset is enriched with confirmed cancers; a real screening population is overwhelmingly people without the disease, where even a 1% false-positive rate generates a flood of downstream investigation of the pancreas — a difficult organ to biopsy safely. The performance a model shows on a curated test set is an upper bound on what it will do in the clinic, not an estimate of it.

What to watch

Whether any pancreatic-detection model is tested prospectively, on consecutive unselected scans, with follow-up long enough to count the cancers found early and the false alarms generated. Whether performance holds on scanners and populations outside the centres that built these models. And whether the downstream pathway — what happens to a patient the AI flags — is worked out before such a tool is switched on, because a detection with no safe, proportionate next step is not yet a benefit.

Sources

  1. [s1] Large-scale pancreatic cancer detection via non-contrast CT and deep learning. Nature Medicine, 20 November 2023. https://doi.org/10.1038/s41591-023-02640-w
  2. [s2] Radiomics for early detection of pancreatic cancer: a systematic review and meta-analysis. Journal of Gastrointestinal Surgery, 16 February 2026. https://doi.org/10.1016/j.gassur.2026.102374

Sources

  1. Large-scale pancreatic cancer detection via non-contrast CT and deep learning — Nature Medicine , November 20, 2023
  2. Radiomics for early detection of pancreatic cancer: a systematic review and meta-analysis — Journal of Gastrointestinal Surgery , February 16, 2026

More on

Related coverage