https://onlinelibrary.wiley.com/doi/epdf/10.1111/acel.70635
This is a really useful approach to research on methylation as a metric.
chatGPT(5.6paid):
Overall assessment
This is a useful and timely methodological paper. Its central message is persuasive: an epigenetic clock can be analytically precise when the same DNA sample is assayed twice, yet fluctuate substantially when blood is collected twice from the same person under ordinary physiological conditions.
The paper’s strongest contribution is distinguishing:
- Technical reliability — whether repeat laboratory assays of the same biological sample agree.
- Biological reliability — whether separate blood samples from the same person give similar results over hours or days.
However, the authors sometimes treat all short-term biological variation as undesirable noise. Their analyses cannot reliably distinguish measurement instability from genuine, reversible biological responses. Several stronger claims—especially about immune-cell adjustment and clinical trustworthiness—therefore go beyond what the data establish.
Summary
Sehgal and colleagues evaluate 18 DNA-methylation ageing biomarkers spanning several clock generations, including:
- Horvath and Hannum chronological-age clocks
- PhenoAge and GrimAge
- GrimAge version 2
- DunedinPACE
- principal-component versions such as PCGrimAge and PCPhenoAge
- SystemsAge, OMICmAge and CausAge/CausalAge
- DNAmEMRAge
They use their TranslAGE pipeline to calculate all clocks consistently across several public datasets.
1. Technical reliability is generally high
Technical reliability was tested using repeat assays of the same DNA, including:
- 196 people with duplicate EPIC-array measurements
- 128 people with duplicate 450K-array measurements
- four people with 16 replicates each involving slide-position variation
- ten people tested using three DNA-extraction methods
Most clocks had pooled technical intraclass correlation coefficients above 0.9. PC-derived clocks, PCGrimAge and SystemsAge were among the strongest performers.
Nevertheless, reliability deteriorated under some specific technical perturbations:
- Different slide positions reduced reliability for a number of clocks.
- Different DNA-extraction procedures affected a smaller subset.
- Earlier clocks such as Hannum and PhenoAge were generally less technically reliable than the best PC-based measures.
Thus, ordinary within-study replicates look reassuring, but some laboratory workflow differences remain important.
2. Separate blood collections are much less stable
“Biological reliability” was assessed from repeated blood collections during:
- meal consumption: 34 participants, four time points in one day
- acute psychosocial stress: 34 participants, four time points
- high-altitude exposure: 21 participants
- short-term pollution exposure: 16 participants
Pooled biological ICCs were mostly approximately 0.4–0.7, and no clock reached the authors’ “excellent” threshold of 0.9.
The most unstable conditions were generally meals and acute stress. GrimAge, GrimAge V2, DNAmEMRAge and DunedinPACE tended to have relatively low biological reliability. PCGrimAge, PCPhenoAge, SystemsAge and OMICmAge were generally more stable, although still much less stable than their technical-replicate results suggested.
Technical and biological reliability were almost completely uncorrelated across clocks:
[
r=0.0168
]
This is one of the paper’s most important results. Good assay reproducibility does not guarantee that a clock will be stable between separate blood collections.
3. Adjusting for immune-cell composition usually lowers ICC
The authors estimate seven leukocyte fractions from the methylation data and residualise clock values for:
- chronological age alone; or
- chronological age plus estimated cell fractions.
Contrary to their expectation, adjustment for cell composition usually reduced biological ICC. The pooled decline across clocks was highly statistically significant.
They interpret this as evidence that immune-cell composition contributes meaningful biological signal rather than merely adding noise. DunedinPACE and PhenoAge were exceptions in some conditions and sometimes improved after adjustment.
4. Less reliable clocks give less stable downstream results
The authors use repeated selection between technical replicates to examine how clock measurement variation affects:
- associations with subsequent MMSE cognitive scores in ADNI
- estimated responses to an eight-week vegan-diet intervention.
Clocks with lower reliability produced much wider distributions of association statistics and intervention effect estimates. Some estimates changed direction depending on which technical replicate was selected.
They report strong inverse correlations between reliability and downstream variability:
- Prognostic z-score variability: (r=-0.906)
- Intervention-response variability: (r=-0.732)
The practical conclusion is that a nominal association or intervention effect obtained using an unreliable clock may not survive repeat measurement.
What is genuinely novel?
1. Joint comparison of technical and biological reliability
Previous work had already shown that epigenetic clocks can have poor technical reproducibility and that epigenetic age varies with time of day, diet and stress. The important advance here is evaluating both forms of reliability across a common set of 18 clocks and demonstrating that they are essentially independent.
The near-zero correlation is more informative than merely reporting that individual clocks fluctuate.
2. Direct linkage to downstream analytic instability
The paper does more than rank clock ICCs. It shows how replicate selection propagates into variability in:
- disease-prognostic associations
- apparent intervention effects.
Although it is unsurprising statistically that noisy measurements destabilise associations, the empirical demonstration is valuable for ageing-intervention research, where small clock changes are often interpreted as biological rejuvenation.
3. Benchmarking newer clocks under ordinary perturbations
The comparison of PC-based, system-level and mechanistically motivated clocks against older clocks is useful. It suggests that PCGrimAge and SystemsAge combine relatively good technical and biological reliability, whereas DunedinPACE illustrates that excellent technical reliability does not necessarily confer short-term biological stability.
4. The immune-cell-adjustment result
The observation that cell-fraction adjustment normally lowers ICC is interesting and potentially important. It challenges the routine assumption that leukocyte adjustment necessarily improves blood-based methylation biomarkers.
What is novel is the observation; the paper’s biological interpretation of it is less secure.
Critique
1. “Biological unreliability” is not necessarily measurement error
A meal, acute stress and pollution exposure can cause genuine changes in:
- circulating leukocyte proportions
- leukocyte activation states
- hormone signalling
- metabolic state
- chromatin regulation and methylation-related measurements.
A clock that responds to these processes may be functioning exactly as trained, particularly if it incorporates immune or metabolic risk signals.
The study establishes short-term within-person variability, but not that this variability is false or biologically meaningless. Calling it unreliability embeds a normative assumption that biological age should remain essentially fixed across hours or days.
A better distinction would be:
- technical measurement error;
- predictable acute biological responsiveness;
- unexplained within-person variation;
- persistent change in the underlying ageing trajectory.
The present datasets cannot fully separate these components.
2. There are no matched time-of-day control groups
The authors acknowledge this important limitation. Blood samples were collected at different times, but comparable participants were not sampled at the same times without the meal, stress or exposure.
Consequently, an apparent “meal effect” could contain:
- circadian variation
- time since waking
- repeated-phlebotomy effects
- posture or hydration changes
- ordinary within-day immune variation
- the meal itself.
The study therefore demonstrates instability under the overall sampling protocol, not necessarily a causal effect of each named perturbation.
3. ICC is heavily dependent on population heterogeneity
ICC is approximately the proportion of total variance attributable to between-person differences:
[
ICC=\frac{\sigma^2_{\text{between}}}
{\sigma^2_{\text{between}}+\sigma^2_{\text{within}}}
]
It is not a context-free property of a biomarker. Three biological datasets consist mainly of young participants, often with narrow age ranges. Restricted between-person variation mechanically lowers ICC even if the absolute within-person measurement variation is modest.
Conversely, a heterogeneous older or diseased cohort can have a high ICC because between-person differences are large, despite substantial short-term fluctuation.
Therefore:
- ICCs from different cohorts are not straightforwardly comparable.
- Low ICC in young homogeneous participants may overstate the problem for an older heterogeneous clinical population.
- Absolute within-person error, repeatability coefficients and coefficients of variation should have been reported alongside ICC.
Age residualisation may further reduce between-person variance and consequently lower ICC.
4. The interpretation of immune-cell adjustment is too strong
The conclusion that immune-cell composition “maintains” reliability or supplies meaningful ageing signal does not necessarily follow.
Regressing clock scores on cell fractions removes both:
- transient within-person cell-composition changes; and
- stable between-person differences in immune composition.
If stable immune differences account for much of the separation between participants, removing them reduces the numerator of the ICC. The ICC can then fall even if unwanted within-person noise has also been reduced.
Furthermore, the cell fractions are inferred from the same methylation data used to calculate the clocks. This creates statistical dependence between the predictor being removed and the clock score.
A more decisive analysis would decompose immune-cell effects into:
- each participant’s mean cell composition;
- deviation from that participant’s mean at each collection.
That would separate stable between-person immune phenotype from acute within-person cell shifts. Purified-cell or single-cell methylation data would be still more informative.
Thus, the paper shows that conventional residualisation lowers ICC; it does not prove that cell adjustment removes beneficial ageing biology.
5. The pooled ICC method is questionable
The authors Fisher-transform ICCs and use a standard error of:
[
SE=1/\sqrt{n-3}
]
This is the familiar approximation for an ordinary Pearson correlation. ICCs have more complicated sampling distributions that depend on:
- the ICC model
- number of participants
- number of replicates
- variance components
- balanced versus unbalanced measurements.
The datasets also differ markedly in design—from 196 people with two replicates to four people with 16 replicates. A hierarchical variance-components model would have been more defensible than treating the ICCs approximately like ordinary correlations.
The reported (I^2>50%) for most clocks confirms substantial heterogeneity. A single pooled “biological reliability” value may therefore obscure important condition-specific behaviour.
6. The downstream analysis conflates technical and biological reliability
The ADNI analysis varies which technical replicate is selected. Its variability should primarily be governed by technical reliability. Yet the results section repeatedly relates downstream dispersion to “biological ICC.”
This needs clearer justification because the paper’s own principal result is that technical and biological ICCs are uncorrelated. It is difficult to claim that the technical-replicate resampling demonstrates the consequences of biological instability.
The intervention example is particularly weak:
- only eight vegan-intervention participants had duplicate assays;
- effect estimates from such a small subset will themselves be highly unstable;
- the presentation of 1,000 resampling iterations can create an impression of greater evidential sample size, although the underlying biological sample remains eight people;
- resampled analyses are not independent replications.
The analysis convincingly illustrates error propagation, but the exact correlations should not be treated as precise universal estimates.
7. Reliability is necessary but not sufficient
The paper occasionally elevates reliability into the “central determinant” of clinical usefulness. A perfectly stable clock can still be:
- biologically invalid
- poorly calibrated
- insensitive to real improvement
- dominated by irreversible chronological-age information
- unrelated to morbidity or treatment benefit.
PC transformation improves reliability partly by averaging correlated CpGs and suppressing noisy components. But such smoothing could also attenuate a genuine intervention response. Stability and responsiveness must therefore be balanced rather than assuming maximum stability is always preferable.
A biomarker for long-term risk should resist irrelevant acute perturbations. A pharmacodynamic biomarker, however, may be valuable precisely because it responds rapidly to biological change.
8. “Technical reliability is solved” is an overstatement
The strongest technical datasets used older 450K and EPIC v1 arrays. The study did not comprehensively test:
- EPIC v2
- different laboratories
- reagent lots
- long storage intervals
- shipping conditions
- different preprocessing pipelines
- clinical-scale assay deployment.
Slide-position reliability was based on only four people, and extraction-method reliability on ten. The results support “generally good within-study repeatability,” not that technical reliability has been solved.
9. Generalisability is limited
Most biological-reliability participants were young adults, whereas epigenetic clocks are commonly considered for older people with multimorbidity. Ageing, inflammation, clonal haematopoiesis and immune-cell heterogeneity could alter both absolute variability and ICC.
The results also come from several small, differently designed datasets rather than a prospectively designed reliability study with standardised repeated sampling.
10. Conflicts and reproducibility deserve attention
The authors disclose that Sehgal and Higgins-Chen are co-inventors on a SystemsAge patent application, and some authors have consulting relationships with commercial longevity-testing organisations. SystemsAge is repeatedly identified as a leading performer.
This does not invalidate the findings, but it strengthens the case for independent replication. In addition, code for OMICmAge and DNAmEMRAge is not openly available in the same way as the other clocks and must be requested through TruDiagnostic, limiting complete independent reproduction.
Practical implications
For intervention studies using epigenetic clocks, the paper supports:
- collecting baseline and follow-up blood at the same time of day;
- standardising fasting status, recent meals, exercise, stress and posture;
- documenting infections, medications and acute exposures;
- randomising matched samples across slides, plates and batches;
- preferably assaying baseline and follow-up samples together;
- using duplicate assays or averaging replicates when expected changes are small;
- prespecifying cell-composition adjustment rather than choosing it after seeing results;
- reporting absolute change relative to the clock’s short-term repeatability;
- prioritising clocks with demonstrated outcome validity as well as reliability.
Bottom line
The paper makes a strong case that excellent technical reproducibility does not guarantee stable estimates between separate blood draws. This is highly relevant to claims that an intervention has reversed epigenetic age by a small amount.
Its most convincing conclusions are:
- technical and short-term biological repeatability are distinct;
- common physiological and sampling conditions materially affect many clocks;
- replicate variation can destabilise prognostic and intervention estimates;
- PCGrimAge and SystemsAge appear comparatively robust in these datasets.
Its less convincing conclusions are that short-term variation is necessarily meaningless ageing noise, that reduced ICC after cell adjustment proves immune composition is beneficial signal, and that reliability alone determines clinical trustworthiness. The paper identifies an important measurement problem, but does not yet establish which fluctuations are artefacts and which are genuine components of dynamic biological ageing.