WHAT THE STUDY ACTUALLY SAYS

A model screened 2.6 million cancer papers and flagged nearly one in ten

Trained on 2,202 retracted paper mill articles, a text classifier flagged 9.87 percent of the cancer literature — including in the highest-impact journals. Flagged is not the same as fraudulent.

Paper mills sell authorship on fabricated scientific papers. They are known to exist, known to be growing, and — until now — measured mostly by counting the ones that got caught, which is a measure of detection rather than prevalence.

A study published in The BMJ on 29 January attempts a prevalence estimate instead, by training a machine learning model on papers already known to have come from mills and then running it across the cancer literature [s1]. The result is large enough that how it is read matters as much as what it says.

The method

The researchers applied a BERT-based text classification model to article titles and abstracts [s1]. Training data came from retracted paper mill publications listed in the Retraction Watch database — 2,202 papers — and the model was validated on independent data collected by image integrity experts [s1].

They then screened PubMed, restricted to original cancer research articles published between 1999 and 2024 [s1]. That corpus contained 2,647,471 papers [s1].

The model achieved an accuracy of 0.91 [s1].

What it found

Applied to the corpus, the model flagged 261,245 papers — 9.87% (95% CI 9.83 to 9.90) [s1].

The distribution over time shows a large increase in flagged papers from 1999 to 2024, both across the entire corpus and within the top 10% of journals by impact factor [s1]. That second point is the one publishers will find hardest. The comfortable assumption has been that mill output collects in low-impact venues; this analysis says otherwise.

More than 170,000 flagged papers were affiliated with Chinese institutions, accounting for 36% of Chinese cancer research articles [s1]. Most publishers had published substantial numbers of flagged papers [s1]. Flagged papers were overrepresented in fundamental research and in gastric, bone and liver cancer [s1].

The authors' conclusion is that paper mills are a large and growing problem in the cancer literature, are not restricted to low-impact journals, and that collective awareness and action will be crucial [s1].

What "flagged" means, and what it does not

This is the part that should govern how the 9.87% figure is used.

The model was trained to identify textual similarity to retracted paper mill publications. The paper's own framing is precise: it assesses the prevalence of papers that have textual similarities to paper mill papers [s1]. It does not determine that any individual flagged paper is fraudulent, and it cannot.

Two failure modes follow directly. Paper mills produce formulaic prose, but so do genuine researchers writing in a second language under templated journal requirements — and the training set's heavy weighting toward one region's retractions could cause the model to learn regional writing conventions alongside fraud signals. A 36% flag rate for one country's cancer output is consistent with a serious integrity problem; it is also consistent with a classifier that has partly learned to detect a writing style.

An accuracy of 0.91 against retracted papers also does not translate simply. Retracted mill papers are, by definition, the ones that were detected — and detection has historically favoured the crudest examples. A model trained on those may be well calibrated for obvious mills and poorly calibrated for careful ones.

None of this means the finding is wrong. It means the honest reading is that roughly one in ten cancer papers resembles known fraudulent output closely enough to warrant a human look, which is itself an alarming amount of human looking to require.

The retraction record, from a different angle

A study published ten days earlier in Arthritis Care & Research takes the conventional approach — counting retractions — in a single specialty, and gives a useful sense of scale [s2].

Reviewing the Retraction Watch database for rheumatology, the authors identified 381 retracted articles between 1989 and 2024, 79.5% of them original articles [s2]. Most originated in Asia (68.5%), particularly China (50.7%) [s2]. Scientific misconduct accounted for 75.3% of retractions, data errors 14.9%, and other reasons 7.6% [s2]. Common misconduct types were data fabrication, fake peer review, duplication and authorship issues [s2].

The trend is steep: 18 retractions in 2000-2009, 117 in 2010-2019, and 207 in 2020-2023 (p < 0.001) [s2]. Compared with other medical specialties, the authors found rheumatology's retraction patterns similar, differing mainly in geographic distribution [s2].

And the number that connects the two studies: the median time from publication to retraction was 18 months (IQR 9 to 46), with one third of articles taking more than 36 months, and time to retraction did not improve over the study period [s2].

Why the two together are the story

Retraction is the correction mechanism, and it is running at 381 papers over 35 years in one specialty, with a median lag of a year and a half that is not getting shorter [s2]. Screening flagged 261,245 papers in one field in one pass [s1].

Even if the screening estimate is several times too high, the mismatch in throughput is the structural problem. The correction mechanism was built for occasional misconduct and is being asked to handle industrial output.

The rheumatology authors offer the standard remedies — early-career education, institutional oversight, ethical research culture [s2]. Those are reasonable and slow. The screening study offers a faster instrument whose false positive behaviour is not yet characterised, which is why its authors frame the output as papers requiring attention rather than papers to be withdrawn [s1].

What to watch

Whether any publisher applies a classifier of this kind prospectively at submission, and whether the false positive rate is published when they do — because a 9.87% flag rate applied to individual careers without human adjudication would create a second integrity problem alongside the first.

Sources

  1. [s1] Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional study. BMJ, 29 January 2026. https://doi.org/10.1136/bmj-2025-087581
  2. [s2] Retractions in Rheumatology: Trends, Causes, and Implications for Research Integrity. Arthritis Care & Research, 19 January 2026. https://doi.org/10.1002/acr.80005

Sources

  1. Machine learning based screening of potential paper mill publications in cancer research: methodological and cross sectional studyBMJ , January 29, 2026
  2. Retractions in Rheumatology: Trends, Causes, and Implications for Research IntegrityArthritis Care & Research , January 19, 2026
Related coverage