ANALYSIS

Run the same sample through an ageing clock twice and it can differ by years

Epigenetic clocks estimate a 'biological age' from DNA methylation. A 2022 study found technical noise alone could shift the result by up to 9 years between replicates of the same sample.

Before asking whether an epigenetic "ageing clock" measures anything meaningful, there is a plainer question: does it give the same answer twice? For the best-known clocks, the answer has been no. A 2022 study in Nature Aging found that technical noise alone could shift a clock's estimate by up to 9 years between two replicates of the same DNA sample [s1] — a reliability problem that sits underneath every claim about what a biological-age number means or whether an intervention moved it.

What a clock is

An epigenetic clock is a statistical model that reads DNA methylation — chemical marks on DNA that change with age — and outputs an estimated age. The original and most cited example, published by Steve Horvath in 2013, was built from 8,000 samples spanning 51 healthy tissues and cell types and distilled the estimate down to 353 specific methylation sites that together form the clock [s2]. It behaved impressively: the estimate sat near zero for embryonic and stem cells, and across 20 cancer types the tissue looked an average of 36 years "older" than the patient [s2]. That performance is why the clocks became the leading candidate biomarker for trials of anti-ageing interventions.

The reliability problem

The trouble is what happens when you measure the same person twice. Methylation is read on microarrays, and the raw signal at any single site carries measurement error. Because a clock sums across many sites, that noise can accumulate into a large swing in the final age estimate even when nothing about the person has changed.

The 2022 study quantified it directly. For six prominent epigenetic clocks, the deviation between replicate measurements of the same sample reached up to 9 years [s1]. A gap of that size is disqualifying for the uses people most want: if a clinic reports that your biological age fell three years after a programme, a swing of that magnitude could be noise rather than any real change [s1]. The authors describe the underlying data as "surprisingly unreliable" [s1] — a striking admission about biomarkers already being sold to consumers and written into study designs.

The fix, and its limits

The same paper proposed a solution. Instead of feeding the clock the raw values from individual methylation sites, the authors computed principal components — summary dimensions that pool information across many sites and average out site-level noise — and trained the clocks on those [s1]. The retrained "principal-component" versions of six clocks brought most replicates into agreement within about 1.5 years, and improved the detection of both clock associations and intervention effects, with more reliable trajectories over time [s1]. The method adds only one step to the standard calculation and needs no repeated samples to train [s1].

Two things follow. First, the problem is fixable, which is genuine good news for the field. Second, and more relevant to a reader, most of the clocks in wide circulation — and behind many consumer tests and older published studies — are the earlier, noisier versions. A result generated by a first-generation clock inherits the reliability limits this study documented, whatever the accompanying marketing says.

Why this matters more than it sounds

Reliability is the unglamorous property that everything else depends on. A biomarker can only track a real change if the real change is larger than the measurement's own wobble. When the wobble is several years, small reported improvements — the kind longevity products advertise — cannot be distinguished from noise without careful design, repeat sampling and the more reliable clock versions.

This is a separate issue from validity, the question of whether the number reflects true biological ageing at all, which the broader evidence on biological-age tests treats sceptically. It also helps explain why some well-run studies find clocks fail to move when they "should": a noisy instrument produces null and spurious results alike. A clock has to clear the reliability bar first; only then does its reading become worth interpreting.

How to read a clock result

Three questions separate a meaningful clock result from a marketed one. Which clock was used, and was it a reliability-corrected (principal-component) version or a first-generation one [s1]? Was the change larger than the several-year measurement error the older clocks carry [s1]? And was the sample measured more than once, so that a reported shift is distinguishable from technical noise? A biological-age number delivered without those safeguards is not necessarily wrong — but on current evidence it is not precise enough to build a personal decision on.

This article is informational and is not medical advice.

Sources

  1. [s1] A computational solution for bolstering reliability of epigenetic clocks: implications for clinical trials and longitudinal tracking. Nature Aging, 2022. https://doi.org/10.1038/s43587-022-00248-2
  2. [s2] DNA methylation age of human tissues and cell types. Genome Biology, 2013. https://doi.org/10.1186/gb-2013-14-10-r115

Sources

  1. A computational solution for bolstering reliability of epigenetic clocks: implications for clinical trials and longitudinal trackingNature Aging , July 15, 2022
  2. DNA methylation age of human tissues and cell typesGenome Biology , October 20, 2013
Related coverage