ANALYSIS

95% of human-microbiome-into-mice studies reported the disease transferred too

A systematic review in Cell called that rate implausible and said it overstates the gut microbiome's role in human disease. A separate audit found the statistical tools themselves disagree.

Human-microbiota-associated rodent studies, and how many reported the phenotype transferredStudies reviewed: 38 studies; Reported phenotype transfer: 36 studies0 studies20 studies40 studiesStudies reviewed38 studiesReported phenotype transfer36 studies
Human-microbiota-associated rodent studies, and how many reported the phenotype transferred
GroupValue (studies)
Studies reviewed38
Reported phenotype transfer36
Human-microbiota-associated rodent studies, and how many reported the phenotype transferred Systematic review of published studies transferring human gut microbiota into rodents. The review reports 95% (36/38) transferring a pathological phenotype. Source: Cell

The central move in modern microbiome science is to take stool from a person with a disease, put it into a germ-free mouse, and see whether the mouse gets sick. When researchers systematically reviewed the published studies that had done this, 95% of them — 36 out of 38 — reported that the pathological phenotype transferred to the recipient animals [s1]. The reviewers' verdict on that number was blunt: they posited that such an exceedingly high rate of inter-species transferable pathologies is implausible and overstates the role of the gut microbiome in human disease [s1].

That assessment, published in Cell in January 2020, is the most important sentence a reader can carry into any consumer claim about the gut microbiome.

Why the number is the problem

Human microbiota-associated rodents have become, in the review's words, a cornerstone of microbiome science for addressing causal relationships between altered microbiomes and host pathology [s1]. The design is genuinely powerful: if a phenotype travels with the microbiota into a sterile animal, that is real evidence the microbiota is doing something.

But a near-universal success rate across a body of literature is a signature, not a triumph. Biology does not usually cooperate at 95%. When almost every published attempt at a difficult transfer works, the plausible explanations include publication bias, flexible outcome definitions, and small studies with high false-positive rates — not a discovery that human diseases are routinely transmissible to mice through faecal material.

The review's authors note that many of these studies extrapolated their findings to make causal inferences about human diseases [s1]. That is the step where a mouse result becomes a headline, and then a supplement.

Their recommendation is a call for a more rigorous and critical approach to inferring causality, explicitly to avoid false concepts and prevent unrealistic expectations that may undermine the credibility of microbiome science and delay its translation [s1]. The concern is not that the field is worthless. It is that overclaiming now buys a backlash later.

The other layer: the statistics disagree with each other

The rodent-transfer problem sits on top of a more basic one. Before anyone can ask whether a microbial difference causes a disease, someone has to establish that the difference exists — and that step turns out to depend on which software was used.

A study published in Nature Communications compared the performance of 14 differential abundance testing methods across 38 16S rRNA gene datasets, each with two sample groups to compare [s2]. The finding was that these tools identified drastically different numbers and sets of significant sequence variants, and that results depended on how the data had been pre-processed [s2].

Worse, the disagreement was not random. For many tools, the number of features identified correlated with properties of the dataset itself — sample size, sequencing depth, and the effect size of the community differences — rather than with biology alone [s2]. Two methods, ALDEx2 and ANCOM-II, produced the most consistent results across studies and agreed best with the intersection of results from the different approaches [s2]. The authors' recommendation was that researchers use a consensus approach across multiple methods to help ensure robust biological interpretations [s2].

Read together with the rodent review, the picture is of a field where the claim "these bacteria are different in people with X" can vary by analytical choice, and the claim "these bacteria cause X" rests on an animal literature whose own reviewers describe its success rate as implausible.

What this does and does not mean

It does not mean the gut microbiome is irrelevant to health. The reviewers in Cell are microbiome scientists arguing for the field's credibility, not against its subject matter [s1]. Faecal microbiota transplantation for recurrent Clostridioides difficile infection remains the standard demonstration that manipulating gut communities can change a clinical outcome.

What it means is that the gap between a mouse result and a human recommendation is far wider than the way those results are usually reported. A study showing that microbiota from people with a condition induced that condition in mice is, on this evidence, a study whose result was 95% likely in advance of anyone running it [s1].

It also means that the most common shape of consumer microbiome claim — that a particular bacterial genus is low in people with a problem, and that raising it will help — is doing two things the literature does not yet support. It treats a differential-abundance finding as stable when the method that produced it may not be [s2], and it treats an association as causal when the causal evidence rests on a rodent literature its own reviewers have called overstated [s1].

What to watch

The corrective the reviewers ask for is methodological rather than technological: better controls, more critical inference, and less extrapolation from animals to humans [s1]. On the statistical side, the fix is already available and rarely used — running several differential abundance methods and reporting where they agree [s2].

Neither is glamorous, and neither produces a product. But between them they determine whether the next decade of microbiome findings replicate.

Sources

Sources

  1. Establishing or Exaggerating Causality for the Gut Microbiome: Lessons from Human Microbiota-Associated RodentsCell , January 23, 2020
  2. Microbiome differential abundance methods produce different results across 38 datasetsNature Communications , January 17, 2022
Related coverage