chatGPT(6AstraMaxPaid):
The paper makes a persuasive case that predicting chronological age and measuring the rate of brain ageing are different tasks. However, its broad rejection of brain-age measures goes further than its evidence supports.
The uploaded paper is Reconsidering Brain Age: Why Age-Prediction Models Fail as Measures of Brain Aging, by Grodem and colleagues. This version is a medRxiv preprint posted on 8 May 2026.
The central argument concerns the “brain-age gap”: the difference between age predicted from brain imaging and actual chronological age. A positive gap is often interpreted as accelerated ageing.
The authors explain why that interpretation is unreliable. Brain measurements contain several kinds of variation:
- Normal changes associated with chronological age.
- Stable differences between people, established through development.
- Differences in how rapidly individuals change over time.
- Measurement noise.
An algorithm trained to predict chronological age favours measurements that change consistently with age. Differences in individual ageing trajectories can make age prediction less accurate, encouraging the algorithm to downweight precisely the variation an ageing-rate biomarker should detect.
For example, someone might have a relatively small hippocampus throughout adulthood. A model could label their brain as older even if their hippocampus is shrinking at an entirely normal rate.
The authors investigate this using mathematical analysis, a simulation and human imaging data.
Their simulation contains two features: one changes at the same rate in everyone, while the other develops increasing differences between individual trajectories. The more accurate age predictor increasingly relies on the first feature and suppresses the second. This demonstrates a plausible situation in which better age prediction produces a worse measure of individual ageing.
For the human analyses, they train Elastic Net and XGBoost models using 121 structural MRI features from approximately 42,600 UK Biobank participants, then examine two contrasting situations.
First, birth weight appears to affect anatomical size more clearly than subsequent shrinkage.
Among 26,695 participants, higher birth weight was associated with a slightly lower brain-age gap. The partial correlation was about -0.05 for both models. Higher birth weight was also associated with larger hippocampal and total brain volumes.
However, among 1,922 participants with repeated MRI measurements, birth weight was not significantly associated with the rate of volume loss. The authors interpret this as evidence that brain-age models can misclassify lifelong anatomical differences as differences in ageing.
The brain-age association was small: a partial correlation of 0.05 corresponds to approximately 0.25% shared residual variance.
Second, ordinary volume measurements were more strongly associated with tau pathology than the brain-age scores.
The tau analysis included 1,199 participants, of whom 856 had longitudinal MRI data. The reported associations with PET tau positivity were:
| Measurement, oriented towards greater abnormality | Adjusted odds ratio | 95% confidence interval |
|---|---|---|
| Higher Elastic Net brain-age gap | 1.21 | 1.06 to 1.38 |
| Higher XGBoost brain-age gap | 1.23 | 1.06 to 1.42 |
| Smaller hippocampal volume | 1.47 | 1.28 to 1.70 |
| Smaller total brain volume | 1.54 | 1.07 to 2.20 |
| Faster hippocampal volume loss | 2.41 | 1.65 to 3.52 |
| Faster total brain volume loss | 2.97 | 1.51 to 5.85 |
These figures come from Figure 3 and the results on pages 9-10. They describe associations, not diagnostic detection rates.
Even a single hippocampal volume measurement had a significantly stronger association than the Elastic Net brain-age gap, with p = 0.0034.
The authors recommend training models to predict longitudinal change, or using selected regional measurements, rather than optimising chronological-age prediction.
The novelty lies mainly in the formal framework and the combined empirical demonstration.
The underlying warning is already established. Vidal-Pineiro and colleagues reported in 2021 that brain-age differences related more to early-life factors than to longitudinal brain change. The birth-weight argument therefore substantially extends earlier work rather than introducing a new discovery. (PubMed)
Likewise, a 2024 methodological paper examined the assumptions needed to interpret cross-sectional age predictors biologically. A 2025 study showed that less accurate brain-age models could have larger disease-related effects, and discussed the suppression of features with greater individual variability. (Springer Nature Link)
The distinctive contributions here are:
- A mathematical decomposition separating measurement noise, baseline differences and individual changes.
- An explanation of how flexible age predictors can increasingly suppress features whose variability grows through individual ageing.
- A paired empirical test involving developmental anatomy and tau-associated degeneration.
- A direct comparison showing that a simple hippocampal measurement can outperform the composite score for a particular biological association.
I would therefore describe this as a useful methodological extension and synthesis, rather than the first demonstration that brain age can be misleading.
The strongest aspect of the study is its alignment of theory, simulation and empirical tests. It offers a reason for the observed failures, uses both linear and nonlinear models, and compares them with interpretable alternatives. Cross-validation and the inclusion of several cohorts also strengthen the work.
My principal criticisms are:
-
Absence of a significant birth-weight association does not establish absence of an effect.
The discussion claims that the longitudinal null result rules out a genuine ageing effect. That is too strong.
For hippocampal change, the correlation was -0.02, with a confidence interval from -0.08 to 0.04. This remains compatible with a small association, including one comparable in magnitude to the reported brain-age association.
The longitudinal sample was also much smaller than the cross-sectional sample. An equivalence analysis, with a specified threshold for a negligible effect, would support the authors’ claim more directly.
-
The tau findings show reduced association, not complete failure.
Both brain-age models were significantly associated with tau positivity. Calling this a “false-negative” scenario risks suggesting that they detected nothing.
Furthermore, larger odds ratios do not themselves establish greater diagnostic sensitivity. That would require measures such as sensitivity at a specified specificity, discrimination and calibration. The evidence supports weaker association relative to some simple measurements.
-
The mathematical conclusion depends on assumptions.
The nonlinear derivation uses a local Gaussian approximation, approximately constant covariance over the relevant age interval, and a locally flat age distribution. The simulation deliberately creates a clean separation between consistent age change and variable individual change.
These are useful demonstrations of a failure mechanism. They do not establish that every age-trained model must lose all useful ageing information. The result depends on the available features, their correlations and how biological changes relate to normal age trends.
-
Accumulated damage, current ageing rate and future risk are different quantities.
A person could have experienced unusually rapid deterioration previously but now be changing at an average rate. A current longitudinal slope would not capture that history.
The paper correctly challenges the interpretation of a brain-age gap as ongoing accelerated ageing. Its broader language sometimes blurs the distinction between measuring current rate and measuring accumulated abnormality.
-
The rejection of brain-health applications is too sweeping.
A measurement can predict future impairment without directly measuring ageing rate. Stable developmental differences could potentially carry information about reserve or vulnerability.
This study does not establish that brain-age scores lack all prognostic value. That requires direct testing against subsequent clinical outcomes, including whether the score adds information beyond ordinary anatomical measurements.
-
The proposed alternatives still require validation.
Predicting future change is better aligned with the intended question, but the paper does not develop and independently validate a replacement model.
Similarly, increasing variability with age is not uniquely attributable to differences in ageing rates: cohort differences, selective survival and measurement effects could also contribute. These alternatives need their own safeguards against overinterpretation.
For healthspan intervention research, the practical implication is that a reduction in predicted brain age should not, by itself, be treated as evidence of slowed or reversed brain ageing. Stronger evidence would combine repeated imaging, reliable estimates of change, cognitive or functional outcomes, and a suitable comparison group. The same question should be asked of other biological clocks: what was the model trained to predict, and what has its change actually been shown to mean?