Biological Versus Technical Reliability of EpigeneticClocks and Implications for Disease Prognosis andIntervention Response (paper July 2026)

https://onlinelibrary.wiley.com/doi/epdf/10.1111/acel.70635

This is a really useful approach to research on methylation as a metric.

chatGPT(5.6paid):

Overall assessment

This is a useful and timely methodological paper. Its central message is persuasive: an epigenetic clock can be analytically precise when the same DNA sample is assayed twice, yet fluctuate substantially when blood is collected twice from the same person under ordinary physiological conditions.

The paper’s strongest contribution is distinguishing:

  • Technical reliability — whether repeat laboratory assays of the same biological sample agree.
  • Biological reliability — whether separate blood samples from the same person give similar results over hours or days.

However, the authors sometimes treat all short-term biological variation as undesirable noise. Their analyses cannot reliably distinguish measurement instability from genuine, reversible biological responses. Several stronger claims—especially about immune-cell adjustment and clinical trustworthiness—therefore go beyond what the data establish.

Summary

Sehgal and colleagues evaluate 18 DNA-methylation ageing biomarkers spanning several clock generations, including:

  • Horvath and Hannum chronological-age clocks
  • PhenoAge and GrimAge
  • GrimAge version 2
  • DunedinPACE
  • principal-component versions such as PCGrimAge and PCPhenoAge
  • SystemsAge, OMICmAge and CausAge/CausalAge
  • DNAmEMRAge

They use their TranslAGE pipeline to calculate all clocks consistently across several public datasets.

1. Technical reliability is generally high

Technical reliability was tested using repeat assays of the same DNA, including:

  • 196 people with duplicate EPIC-array measurements
  • 128 people with duplicate 450K-array measurements
  • four people with 16 replicates each involving slide-position variation
  • ten people tested using three DNA-extraction methods

Most clocks had pooled technical intraclass correlation coefficients above 0.9. PC-derived clocks, PCGrimAge and SystemsAge were among the strongest performers.

Nevertheless, reliability deteriorated under some specific technical perturbations:

  • Different slide positions reduced reliability for a number of clocks.
  • Different DNA-extraction procedures affected a smaller subset.
  • Earlier clocks such as Hannum and PhenoAge were generally less technically reliable than the best PC-based measures.

Thus, ordinary within-study replicates look reassuring, but some laboratory workflow differences remain important.

2. Separate blood collections are much less stable

“Biological reliability” was assessed from repeated blood collections during:

  • meal consumption: 34 participants, four time points in one day
  • acute psychosocial stress: 34 participants, four time points
  • high-altitude exposure: 21 participants
  • short-term pollution exposure: 16 participants

Pooled biological ICCs were mostly approximately 0.4–0.7, and no clock reached the authors’ “excellent” threshold of 0.9.

The most unstable conditions were generally meals and acute stress. GrimAge, GrimAge V2, DNAmEMRAge and DunedinPACE tended to have relatively low biological reliability. PCGrimAge, PCPhenoAge, SystemsAge and OMICmAge were generally more stable, although still much less stable than their technical-replicate results suggested.

Technical and biological reliability were almost completely uncorrelated across clocks:

[
r=0.0168
]

This is one of the paper’s most important results. Good assay reproducibility does not guarantee that a clock will be stable between separate blood collections.

3. Adjusting for immune-cell composition usually lowers ICC

The authors estimate seven leukocyte fractions from the methylation data and residualise clock values for:

  • chronological age alone; or
  • chronological age plus estimated cell fractions.

Contrary to their expectation, adjustment for cell composition usually reduced biological ICC. The pooled decline across clocks was highly statistically significant.

They interpret this as evidence that immune-cell composition contributes meaningful biological signal rather than merely adding noise. DunedinPACE and PhenoAge were exceptions in some conditions and sometimes improved after adjustment.

4. Less reliable clocks give less stable downstream results

The authors use repeated selection between technical replicates to examine how clock measurement variation affects:

  • associations with subsequent MMSE cognitive scores in ADNI
  • estimated responses to an eight-week vegan-diet intervention.

Clocks with lower reliability produced much wider distributions of association statistics and intervention effect estimates. Some estimates changed direction depending on which technical replicate was selected.

They report strong inverse correlations between reliability and downstream variability:

  • Prognostic z-score variability: (r=-0.906)
  • Intervention-response variability: (r=-0.732)

The practical conclusion is that a nominal association or intervention effect obtained using an unreliable clock may not survive repeat measurement.

What is genuinely novel?

1. Joint comparison of technical and biological reliability

Previous work had already shown that epigenetic clocks can have poor technical reproducibility and that epigenetic age varies with time of day, diet and stress. The important advance here is evaluating both forms of reliability across a common set of 18 clocks and demonstrating that they are essentially independent.

The near-zero correlation is more informative than merely reporting that individual clocks fluctuate.

2. Direct linkage to downstream analytic instability

The paper does more than rank clock ICCs. It shows how replicate selection propagates into variability in:

  • disease-prognostic associations
  • apparent intervention effects.

Although it is unsurprising statistically that noisy measurements destabilise associations, the empirical demonstration is valuable for ageing-intervention research, where small clock changes are often interpreted as biological rejuvenation.

3. Benchmarking newer clocks under ordinary perturbations

The comparison of PC-based, system-level and mechanistically motivated clocks against older clocks is useful. It suggests that PCGrimAge and SystemsAge combine relatively good technical and biological reliability, whereas DunedinPACE illustrates that excellent technical reliability does not necessarily confer short-term biological stability.

4. The immune-cell-adjustment result

The observation that cell-fraction adjustment normally lowers ICC is interesting and potentially important. It challenges the routine assumption that leukocyte adjustment necessarily improves blood-based methylation biomarkers.

What is novel is the observation; the paper’s biological interpretation of it is less secure.

Critique

1. “Biological unreliability” is not necessarily measurement error

A meal, acute stress and pollution exposure can cause genuine changes in:

  • circulating leukocyte proportions
  • leukocyte activation states
  • hormone signalling
  • metabolic state
  • chromatin regulation and methylation-related measurements.

A clock that responds to these processes may be functioning exactly as trained, particularly if it incorporates immune or metabolic risk signals.

The study establishes short-term within-person variability, but not that this variability is false or biologically meaningless. Calling it unreliability embeds a normative assumption that biological age should remain essentially fixed across hours or days.

A better distinction would be:

  1. technical measurement error;
  2. predictable acute biological responsiveness;
  3. unexplained within-person variation;
  4. persistent change in the underlying ageing trajectory.

The present datasets cannot fully separate these components.

2. There are no matched time-of-day control groups

The authors acknowledge this important limitation. Blood samples were collected at different times, but comparable participants were not sampled at the same times without the meal, stress or exposure.

Consequently, an apparent “meal effect” could contain:

  • circadian variation
  • time since waking
  • repeated-phlebotomy effects
  • posture or hydration changes
  • ordinary within-day immune variation
  • the meal itself.

The study therefore demonstrates instability under the overall sampling protocol, not necessarily a causal effect of each named perturbation.

3. ICC is heavily dependent on population heterogeneity

ICC is approximately the proportion of total variance attributable to between-person differences:

[
ICC=\frac{\sigma^2_{\text{between}}}
{\sigma^2_{\text{between}}+\sigma^2_{\text{within}}}
]

It is not a context-free property of a biomarker. Three biological datasets consist mainly of young participants, often with narrow age ranges. Restricted between-person variation mechanically lowers ICC even if the absolute within-person measurement variation is modest.

Conversely, a heterogeneous older or diseased cohort can have a high ICC because between-person differences are large, despite substantial short-term fluctuation.

Therefore:

  • ICCs from different cohorts are not straightforwardly comparable.
  • Low ICC in young homogeneous participants may overstate the problem for an older heterogeneous clinical population.
  • Absolute within-person error, repeatability coefficients and coefficients of variation should have been reported alongside ICC.

Age residualisation may further reduce between-person variance and consequently lower ICC.

4. The interpretation of immune-cell adjustment is too strong

The conclusion that immune-cell composition “maintains” reliability or supplies meaningful ageing signal does not necessarily follow.

Regressing clock scores on cell fractions removes both:

  • transient within-person cell-composition changes; and
  • stable between-person differences in immune composition.

If stable immune differences account for much of the separation between participants, removing them reduces the numerator of the ICC. The ICC can then fall even if unwanted within-person noise has also been reduced.

Furthermore, the cell fractions are inferred from the same methylation data used to calculate the clocks. This creates statistical dependence between the predictor being removed and the clock score.

A more decisive analysis would decompose immune-cell effects into:

  • each participant’s mean cell composition;
  • deviation from that participant’s mean at each collection.

That would separate stable between-person immune phenotype from acute within-person cell shifts. Purified-cell or single-cell methylation data would be still more informative.

Thus, the paper shows that conventional residualisation lowers ICC; it does not prove that cell adjustment removes beneficial ageing biology.

5. The pooled ICC method is questionable

The authors Fisher-transform ICCs and use a standard error of:

[
SE=1/\sqrt{n-3}
]

This is the familiar approximation for an ordinary Pearson correlation. ICCs have more complicated sampling distributions that depend on:

  • the ICC model
  • number of participants
  • number of replicates
  • variance components
  • balanced versus unbalanced measurements.

The datasets also differ markedly in design—from 196 people with two replicates to four people with 16 replicates. A hierarchical variance-components model would have been more defensible than treating the ICCs approximately like ordinary correlations.

The reported (I^2>50%) for most clocks confirms substantial heterogeneity. A single pooled “biological reliability” value may therefore obscure important condition-specific behaviour.

6. The downstream analysis conflates technical and biological reliability

The ADNI analysis varies which technical replicate is selected. Its variability should primarily be governed by technical reliability. Yet the results section repeatedly relates downstream dispersion to “biological ICC.”

This needs clearer justification because the paper’s own principal result is that technical and biological ICCs are uncorrelated. It is difficult to claim that the technical-replicate resampling demonstrates the consequences of biological instability.

The intervention example is particularly weak:

  • only eight vegan-intervention participants had duplicate assays;
  • effect estimates from such a small subset will themselves be highly unstable;
  • the presentation of 1,000 resampling iterations can create an impression of greater evidential sample size, although the underlying biological sample remains eight people;
  • resampled analyses are not independent replications.

The analysis convincingly illustrates error propagation, but the exact correlations should not be treated as precise universal estimates.

7. Reliability is necessary but not sufficient

The paper occasionally elevates reliability into the “central determinant” of clinical usefulness. A perfectly stable clock can still be:

  • biologically invalid
  • poorly calibrated
  • insensitive to real improvement
  • dominated by irreversible chronological-age information
  • unrelated to morbidity or treatment benefit.

PC transformation improves reliability partly by averaging correlated CpGs and suppressing noisy components. But such smoothing could also attenuate a genuine intervention response. Stability and responsiveness must therefore be balanced rather than assuming maximum stability is always preferable.

A biomarker for long-term risk should resist irrelevant acute perturbations. A pharmacodynamic biomarker, however, may be valuable precisely because it responds rapidly to biological change.

8. “Technical reliability is solved” is an overstatement

The strongest technical datasets used older 450K and EPIC v1 arrays. The study did not comprehensively test:

  • EPIC v2
  • different laboratories
  • reagent lots
  • long storage intervals
  • shipping conditions
  • different preprocessing pipelines
  • clinical-scale assay deployment.

Slide-position reliability was based on only four people, and extraction-method reliability on ten. The results support “generally good within-study repeatability,” not that technical reliability has been solved.

9. Generalisability is limited

Most biological-reliability participants were young adults, whereas epigenetic clocks are commonly considered for older people with multimorbidity. Ageing, inflammation, clonal haematopoiesis and immune-cell heterogeneity could alter both absolute variability and ICC.

The results also come from several small, differently designed datasets rather than a prospectively designed reliability study with standardised repeated sampling.

10. Conflicts and reproducibility deserve attention

The authors disclose that Sehgal and Higgins-Chen are co-inventors on a SystemsAge patent application, and some authors have consulting relationships with commercial longevity-testing organisations. SystemsAge is repeatedly identified as a leading performer.

This does not invalidate the findings, but it strengthens the case for independent replication. In addition, code for OMICmAge and DNAmEMRAge is not openly available in the same way as the other clocks and must be requested through TruDiagnostic, limiting complete independent reproduction.

Practical implications

For intervention studies using epigenetic clocks, the paper supports:

  • collecting baseline and follow-up blood at the same time of day;
  • standardising fasting status, recent meals, exercise, stress and posture;
  • documenting infections, medications and acute exposures;
  • randomising matched samples across slides, plates and batches;
  • preferably assaying baseline and follow-up samples together;
  • using duplicate assays or averaging replicates when expected changes are small;
  • prespecifying cell-composition adjustment rather than choosing it after seeing results;
  • reporting absolute change relative to the clock’s short-term repeatability;
  • prioritising clocks with demonstrated outcome validity as well as reliability.

Bottom line

The paper makes a strong case that excellent technical reproducibility does not guarantee stable estimates between separate blood draws. This is highly relevant to claims that an intervention has reversed epigenetic age by a small amount.

Its most convincing conclusions are:

  1. technical and short-term biological repeatability are distinct;
  2. common physiological and sampling conditions materially affect many clocks;
  3. replicate variation can destabilise prognostic and intervention estimates;
  4. PCGrimAge and SystemsAge appear comparatively robust in these datasets.

Its less convincing conclusions are that short-term variation is necessarily meaningless ageing noise, that reduced ICC after cell adjustment proves immune composition is beneficial signal, and that reliability alone determines clinical trustworthiness. The paper identifies an important measurement problem, but does not yet establish which fluctuations are artefacts and which are genuine components of dynamic biological ageing.

Extracting a key part of the paper:

By contrast, biological reliability remains a major obstacle. Wefound that repeated measures taken under common short-termperturbations, including meals, stress, and pollution exposure,caused substantial fluctuations in epigenetic age estimates,with most clocks falling into only moderate reliability at best(Higgins- Chen etal. 2022; Komaki et al. 2022; Galkin et al. 2023).Importantly, these are the types of perturbations that would occur during the course of everyday life, making them commonconfounding events. Adjusting for immune cell composition, wesaw that biological reliability dropped significantly. Critically,technical reproducibility did not predict biological stability:some of the most technically robust clocks, such as GrimAgeV2and DunedinPACE, were among the most biologically fragile.This dissociation reveals that laboratory precision alone is notsufficient to ensure a biomarker’s clinical utility

Another question:

What does this mean in terms of the variation in DunedinPACE that might be expected base upon eating a meal or the time of day

The paper shows that a single DunedinPACE result could vary substantially during the day, but it does not report enough information to convert the meal ICC directly into a precise DunedinPACE range.

What the meal ICC means

For the meal/time-of-day series, DunedinPACE had an ICC of approximately 0.12. In simplified terms:

[
ICC=\frac{\text{between-person variance}}
{\text{between-person variance}+\text{within-person variance}}
]

An ICC of 0.12 implies:

[
\frac{\text{within-person variance}}
{\text{between-person variance}}
=\frac{1-0.12}{0.12}\approx7.3
]

Thus, in this young and relatively homogeneous cohort, short-term variation within the same person was about seven times the between-person variance. Only about 12% of the observed variance represented stable differences between individuals.

That is very poor repeatability.

It was not a consistent “meal effect”

Importantly, the authors did not find that eating systematically increased or decreased DunedinPACE. Instead, the measurements fluctuated inconsistently.

The protocol involved blood collection at approximately:

  • 11:30
  • 12:45
  • 13:45
  • 16:15

with the meal at 14:00.

Consequently, the low ICC represents the combined effects of:

  • time of day;
  • fasting duration;
  • eating a meal;
  • postprandial immune-cell redistribution;
  • hydration and metabolic changes;
  • ordinary measurement-to-measurement biological variation.

There was no control group sampled at the same times without a meal. It is therefore impossible to say how much was caused by the meal itself and how much was circadian or procedural.

Possible variation in DunedinPACE units

DunedinPACE is centred approximately around 1.0:

  • 1.00 means an estimated pace of one biological year per chronological year;
  • 0.90 means approximately 10% slower;
  • 1.10 means approximately 10% faster.

The ICC alone does not reveal the absolute within-person standard deviation. That requires the raw DunedinPACE distribution or the supplementary variance components.

For illustration, if the total standard deviation in a particular cohort were:

Total DunedinPACE SD Approximate SD of difference between two daily measurements at ICC 0.12 Approximate 95% repeatability interval
0.05 0.066 ±0.13
0.08 0.106 ±0.21
0.10 0.133 ±0.26

These are illustrations based on assumed cohort standard deviations, not values directly reported by the paper. They show why the ICC matters: an apparent change such as 0.05—or potentially even 0.10—might be compatible with ordinary within-day variation.

Practical interpretation

A DunedinPACE change from, for example, 0.95 to 1.02 after an intervention should not automatically be interpreted as a genuine seven-percentage-point acceleration of ageing if:

  • one sample was fasting and the other fed;
  • samples were collected at different times;
  • recent stress or exercise differed;
  • the person had an infection or inflammatory event;
  • only one blood draw was taken at each assessment.

The paper’s results imply that DunedinPACE should ideally be measured:

  • at the same morning time;
  • after the same fasting period;
  • before food and strenuous exercise;
  • under similar medication and hydration conditions;
  • using two or more blood collections at baseline and follow-up.

The key conclusion is therefore not that a meal changes the true pace of ageing. It is that a single DunedinPACE measurement taken after a meal or at a different time of day may produce an apparent change large enough to obscure the relatively small effects expected from many ageing interventions.