Source notes
Discrepancies we have found inside published papers and agency bulletins, and what we did about each one.
Why this page exists
Our house rule is that every numeral in an article must appear in the source it is credited to. Checking that, one figure at a time, turns up a steady trickle of errors in the sources themselves — arithmetic that does not close, confidence intervals that do not contain their own point estimates, agency bulletins that disagree with themselves.
None of these are our corrections; our own are on the corrections page. These are problems in material we cited. We publish them because the alternative is silently choosing one figure and presenting it as settled, which would make our reporting look tidier than the evidence is.
How we handle one
We report both figures and attribute each, or we cite the raw counts and omit the derived percentage. We do not pick the more convenient number, and we do not quietly drop the story. Where a discrepancy is large enough that the finding cannot be characterised at all, we do not write the article.
A discrepancy noted here is not an allegation of misconduct. Most are transcription or rounding errors of the kind that survive peer review routinely, and pointing at one is not the same as doubting a paper’s conclusion.
The record
A confidence interval that does not contain its own point estimate
A study of wildfire smoke and pregnancy outcomes reported a hazard ratio of 1.267 for non-movers, with a 95% confidence interval of 1.054 to 1.205. The interval cannot bracket the estimate it belongs to, so at least one of the three numbers is wrong.
What we did: Reported the discrepancy without resolving it in either direction, since the abstract gives no basis to choose.
A percentage that contradicts its own fraction
A study of blood biomarkers in dementia diagnosis stated that a pre-biomarker diagnosis was maintained in "71/200 cases (75.5%)". 71 of 200 is 35.5%. 151 of 200 is 75.5%.
What we did: Reported both readings and flagged that the source does not settle which is intended.
A prose summary that contradicts its own figure
A survey of loneliness among UK university students reported moderate to severe loneliness in 78.98% of 1,408 respondents, while describing that result in the same abstract as "approximately two thirds".
What we did: Quoted the figure, noted the mismatch with the paper's own characterisation, and used the number rather than the description.
An agency bulletin that disagrees with itself
A WHO Disease Outbreak News item on measles in Bangladesh gave 2,897 laboratory-confirmed cases in its summary and 2,973 in its description, for the same reporting period.
What we did: Reported both figures and attributed each to the part of the bulletin it came from.
Two agency updates whose totals do not reconcile
WHO's April MERS update gave 2,627 cases and 946 deaths. Its December update gave 2,635 and 964 — eight more cases but eighteen more deaths, across a period the December bulletin itself describes as nine cases and two deaths.
What we did: Reported the gap alongside WHO's own note about delayed retrospective reporting, rather than presenting either total as settled.
Two reviews, one campaign, different denominators
Two reviews of Egypt's national hepatitis C programme gave different screening totals for the same campaign — 65 million people in one, 57 million in the other.
What we did: Stated both and attributed each, noting the difference sits in the screening denominator rather than the treatment count.
A rounding discrepancy in a headline prevalence
A cohort study of CTE among former NFL players reported a minimum prevalence at death of 18.5%. The underlying counts, 315 of 1,712, give 18.4%. The upper bound in the same paper reproduces exactly.
What we did: Reported the paper's figure and the arithmetic, and characterised it as rounding-level rather than substantive.
Digit separators that do not match their own digit groups
A global survey of PFAS in landfill leachate printed the upper end of its total-PFAS range as "5542,000 ng/L" and its industrial-landfill average as "1141,345 ng/L". Read literally those are 5,542,000 and 1,141,345 ng/L, but the separator placement does not match the grouping in either case.
What we did: Reported both values exactly as printed, flagged the formatting, and did not silently normalise them.
A published abstract and its author manuscript give different sample descriptions
The Science paper measuring human water turnover by isotope tracking describes its 5,604 participants as drawn from 23 countries in the published abstract and from 26 countries in the author manuscript deposited in PubMed Central. The participant count and the age range, 8 days to 96 years, are identical in both versions.
What we did: Reported both figures rather than choosing between them, and used only the counts that agree.
An abstract that gives two sample sizes in consecutive sentences
A chromatography study of commercial melatonin supplements states that melatonin was quantified in 30 commercial supplements, and in the next sentence that a total of 31 supplements were analyzed. The serotonin result — eight products, described as an additional 26% — is consistent with a denominator of 31, but the abstract never reconciles the two counts.
What we did: Reported both sample sizes, showed which of the paper's own figures is consistent with which denominator, and told readers to treat the exact denominator as unresolved.
Two effect sizes five times apart, sharing one confidence interval
An umbrella review of physical therapies for delayed-onset muscle soreness reports effect sizes ranging from a Hedges' g of 0.36 for cooling therapy to 1.82 for heat therapy, and prints the identical 95% confidence interval of 0.46 to 3.18 against both.
What we did: Reported both figures as printed, flagged that one interval must be a transcription error, and told readers not to rely on the precision of either.
A concluding sentence that contradicts the paper's own result
A meta-analysis of detraining and maximal oxygen uptake found a larger decline after long-term training cessation than short-term, then summarised a subgroup comparison of two long-cessation bands as showing no significant change in VO2max beyond 30 days of cessation — which reads as the opposite finding.
What we did: Reported the subgroup result as the paper computed it and flagged that the concluding sentence misstates what was compared.
A significant hazard ratio described as not significant
A UK Biobank analysis of weekend-warrior activity in people with hypertension reports a hazard ratio of 0.59 (95% CI 0.41 to 0.83) with P = 4.0 x 10-3 for cardiometabolic multimorbidity, then states in the same abstract that the result showed no statistical significance.
What we did: Reported both statements, noted that a false-discovery-rate correction is the likely explanation the abstract never gives, and treated the finding as unresolved.
A training volume that cannot fit in a week
A randomised trial of energy compensation reports its six-sessions-a-week group exercising 320.5 minutes per week and its two-sessions-a-week group 1,888.8 minutes per week, in a sentence stating that the six-day group exercised longer.
What we did: Reported the energy figures, which are internally consistent, and flagged the two-day minutes figure as an error that cannot be reconstructed from the text.
A confidence interval crossing zero inside a claim of superiority
A reanalysis of exercise trials for chronic low back pain reports a 2004 disability effect size of -6.67 with a 95% confidence interval of -11.27 to 3.36, in a sentence stating that the superiority boundary had already been crossed.
What we did: Reported the interval as printed, noted that the upper bound is most likely a sign error, and showed that the paper's substantive finding does not depend on it.
A biological-age table whose rows do not reconcile with its own summary
A review of off-label rapamycin modelled PhenoAge in a 25-person trial. Its table lists the rapamycin group's post-treatment phenotypic age as 77.38 against a chronological age of 80.4 — a gap of -3.02 years — while recording -3.96 on that line, which is instead the change in the gap. The control column lists gaps of -2.28 and -1.93, differing by 0.35, while the row beneath states the control difference as 0.15. The control group's chronological age also falls from 80.6 to 80.4 across the study.
What we did: Reported the paper's stated figures and the arithmetic that does not follow from them, alongside the authors' own caveat that significance could not be determined and two inputs were imputed.
A blood concentration reported in units its own thresholds contradict
The same review gives the mean circulating sirolimus level achieved in an eight-week trial as 7.2 ng/dL, while stating elsewhere that the drug has biological effects at about 5 ng/mL and greater toxicity above 15 ng/mL. The two scales differ by a factor of a hundred, and the stated level would fall far below the stated threshold.
What we did: Quoted the figure exactly as printed and flagged the unit mismatch rather than silently converting it.
An abstract crediting a drug with an improvement its placebo arm also showed
The PEARL trial's abstract reports that self-reported emotional well-being improved for participants using 5 mg rapamycin. Its results section shows the improvement reaching significance in the 5 mg group (mean difference 5.176) and in the placebo group (4.267) alike.
What we did: Reported both arms, and noted that the abstract alone supports a stronger reading than the results section does.
A table and a discussion section counting the same replicates differently
A NIST evaluation of consumer gut-microbiome tests gives one company's three replicate samples as 96, 98 and 99 species in Table 1, and as 95, 98 and 99 in its discussion of the same replicates.
What we did: Reported the table's figures and noted the discrepancy, since neither set is identifiable as the correction.
A conclusion that sits on the wrong side of the threshold it names
A meta-analysis of enamel microhardness after peroxide bleaching reported a pooled ratio of means of 0.89 (95% CI 0.84 to 0.94), then concluded there was no clear evidence of clinically meaningful effects, defining those as a ratio of means exceeding 10%. The point estimate is a larger reduction than that threshold, and the confidence interval spans both sides of it.
What we did: Reported the pooled estimate, the interval and the authors' stated threshold, and flagged the tension rather than adopting either reading.
A confidence-interval bound printed an order of magnitude off
The NSABP B-59 trial of atezolizumab in triple-negative breast cancer prints the lower bound of its event-free-survival hazard-ratio interval as 0.062 — inconsistent with its own hazard ratio of 0.80, its P value of 0.083, and its stated non-significant result. It reads as a typo for 0.62.
What we did: Cited the hazard ratio, P value and the paper's own conclusion, and did not reproduce the implausible bound.
A case-fatality rate stated two ways that do not agree
A Yemen diphtheria surveillance registry reports an overall case-fatality rate of 8.04% (44 of 547 cases) and then, describing the same deaths among unvaccinated patients, gives 8.7% — a figure the raw counts do not reconcile.
What we did: Cited the raw counts and the headline case-fatality rate, and did not adopt the inconsistent parenthetical figure.
Telling us we are wrong about one of these
If we have misread a paper, we would rather know. Write to us via our contact page with the article and the specific figure, and the correction will appear on the corrections page.