WHAT THE STUDY ACTUALLY SAYS

Seven of 25 sports science studies replicated on all three criteria

The first large replication project in the field found effect sizes shrank substantially on repeat, and reports that many original authors declined to engage with the attempt.

Psychology has had a replication crisis. So, in different forms, have cancer biology and economics. Sports and exercise science has had the concerns without the audit: as the authors of a paper published in Sports Medicine on 16 June put it, the replicability of the field had not been assessed previously, despite concerns about scientific practices within it [s1].

It has now, and the results are published alongside an unusually candid companion paper about how difficult the exercise was [s1][s2].

The design

The project aimed to produce an initial estimate of the replicability of applied sports and exercise science research published in quartile 1 journals — ranked by the SCImago journal ranking for 2019 in the Sports Science subject category — between 2016 and 2021 [s1].

The selection protocol was published in advance [s1]. Voluntary collaborators were recruited, and studies were allocated in a stratified and randomised way according to the equipment and expertise each team had [s1]. Original authors were contacted to supply de-identified raw data, review preregistrations and clarify methods [s1].

The analysis used a multiple inferential strategy. The same test as the original — an F test or a t test — determined whether the replication effect was statistically significant and in the same direction. Separately, Z tests determined whether the original and replication effect size estimates were compatible or significantly different in magnitude [s1].

Twenty-five replication studies were included: 10 used paired t tests, one an independent t test and 14 an analysis of variance [s1].

The result

Seven of the 25 — 28% — demonstrated what the authors call robust replicability, meeting all three validation criteria: statistical significance at P<0.05, the same direction as the original, and compatible effect size magnitude by the Z test (P>0.05) [s1].

The headline percentage is easy to over-read. Twenty-five studies is a small sample from which to characterise a whole discipline, and the authors present the figure as an initial estimate rather than a verdict [s1]. Failing to replicate a single study does not establish that the original was wrong; it establishes that a second attempt under a different team, sample and setting did not produce a compatible result.

The finding the authors themselves foreground is not the pass rate but the magnitudes: there was a substantial decrease in published effect size estimates when replicated, and they draw the practical conclusion that researchers should account for effect size uncertainty when running subsequent power analyses [s1].

That is a specific and consequential point. Sample sizes in this field are routinely calculated from published effect sizes. If those published estimates are systematically inflated, then studies powered from them are systematically underpowered — which produces more inflated estimates. The loop is self-sustaining.

The part that is usually left out

The companion review, published the same day, describes what the project ran into [s2].

The authors report that preparing studies for replication was obstructed by poor reporting of statistical information and by the availability of original raw data, and that feasibility was prioritised at the risk of some bias [s2]. They state their view that these issues reflect the wider field rather than the particular studies selected [s2].

The sentence that will be quoted is about people rather than data: discourse with original study authors was a challenging process, as many were unwilling to engage, which the authors read as indicating a problematic perception of replication [s2].

They argue that research culture needs to change to minimise active engagement in behaviours that reduce reproducibility and replicability [s2]. That is a strong claim, and it is an argument rather than a measurement — the companion paper is a reflective review, not a study, and should be read as the project team's own account of its experience.

Why this matters outside the field

Sports and exercise science produces a large share of the findings that reach the public as practical advice: this training protocol, that recovery method, this supplement timing. Those findings often rest on small crossover trials with within-subject designs — exactly the paired t-test and ANOVA studies that made up the replication set [s1].

A 28% robust replication rate does not mean 72% of such findings are false. It means that in this sample, most repeat attempts did not reproduce both the significance and the magnitude of the original — and that the magnitudes, in particular, came down.

For a reader, the operational implication is narrow: the size of an effect reported once in a small study is the least reliable number in the paper, and the one most likely to shrink.

What to watch

Whether journals in the field move on the two specific barriers the project named — statistical reporting completeness and raw data availability [s2]. Both are addressable by editorial policy rather than by culture change, and both are measurable.

Sources

Sources

  1. Estimating the Replicability of Sports and Exercise Science ResearchSports Medicine , June 16, 2025
  2. Reflections on Conducting a Large Replication Project in Sports and Exercise ScienceSports Medicine , June 16, 2025
Related coverage