Comparing Biological Age Services: The Questions That Matter
Providers change quarterly and the questions do not. Seven of them separate a defensible service from an expensive one, and most fail at least two.
The Short Answer
The consumer biological age market turns over faster than any comparison can track: services launch, rebrand, switch laboratories and change which clock they report, sometimes without announcing it. A brand ranking is stale within two quarters. The questions that separate a defensible service from an expensive one are stable, and this article is those questions. The underlying science of what the clocks measure is covered separately in the clock generations article.
Question One: Which Clock, and Which Generation
The most informative question, and a surprising number of services will not answer it plainly.
First generation clocks such as Horvath and Hannum were trained to predict chronological age. Accurate at that, comparatively weak at predicting health outcomes.
Second generation clocks such as PhenoAge and GrimAge were trained on clinical markers and mortality, and predict outcomes considerably better.
Third generation, principally DunedinPACE, estimates rate of ageing rather than accumulated age, which is the most decision-relevant of the three.
Proprietary clocks are the category to scrutinise, since an unpublished model asks you to accept validation you cannot inspect. Some are methodologically sound and there is no way for a consumer to tell which.
A good answer names the published clock and cites the paper. A poor answer is a proprietary algorithm based on the latest research, which is a description of nothing.
Question Two: Principal-Component Version or Not
This determines whether a repeat measurement means anything.
Standard methylation clocks read individual CpG sites, and individual site measurement carries real technical variance. Test-retest on the same blood sample can differ by two to three years, which is larger than most changes anyone hopes to detect.
Principal-component versions aggregate across many correlated sites to reduce that noise substantially, and they are the current methodological standard for anything intended to be repeated.
What to ask: whether the service uses a principal-component version, and what its test-retest reliability is. A service that has measured reliability will tell you; one that has not is reporting a number with an uncertainty band it is not showing you.
Why this matters commercially: services encourage repeat testing, and repeat testing on a non-principal-component clock at short intervals largely measures the assay. That is a structural problem rather than an accident.
Questions Three to Five: Sample, Laboratory, Uncertainty
| Question | Good answer | Concerning answer |
|---|---|---|
| What is sampled? | Venous blood, with cell-composition adjustment stated | Saliva with a blood-calibrated clock, or no mention of cell composition |
| Which laboratory? | Named, accredited, and the same across repeats | Unnamed, or variable |
| Is uncertainty reported? | An interval alongside the estimate | A single number implying precision the method lacks |
| Is batch correction applied? | Yes, stated | Not mentioned |
Cell composition is the recurring confounder. A blood sample is a mixture whose proportions shift with recent infection, stress and time of day, and part of the apparent methylation signal is a change in which cells were present. Services that adjust statistically for estimated cell counts and say so are doing the work.
The saliva point matters practically. Saliva is the most convenient sample and its cell composition varies most, and clocks calibrated on blood do not transfer cleanly. If you use saliva, stay with it rather than mixing sample types, since a change of sample produces a change of number.
Uncertainty reporting is the clearest honesty signal. A report stating an estimate plus or minus three years is being more accurate than one stating a single figure.
Questions Six and Seven: Incentives and Data
Six: what does the result route to? If the report arrives with a supplement recommendation from the same company, the interpretation layer is not independent of the sales layer. That does not invalidate the assay, and it means the recommendation should be weighted accordingly. It also creates a structural pressure toward findings that require a purchase.
Seven: what happens to the data? This is the question most worth asking and least asked. Methylation data are among the most identifying and inferentially rich biological data a consumer can generate, carrying information well beyond an age estimate, including markers relevant to smoking history, alcohol exposure and several health conditions.
What to read: retention period, research consent terms, third-party sharing, whether raw data are returned to you, and the deletion path. In most jurisdictions this sits outside the protections that apply to clinical records.
Also worth checking: whether you receive the raw methylation data or only a derived score. Raw data are portable and re-analysable as clocks improve; a score is neither.
A service failing question seven is asking for something more valuable than the fee.
What No Service Can Provide
Being explicit about this prevents the most common disappointment.
Individual decision guidance. No epigenetic clock is validated to indicate which change a specific person should make, and using a result as a target risks optimising the assay rather than the biology.
A meaningful short-interval change. Test-retest variability commonly exceeds the biological change achievable in months, so a three-month follow-up largely measures the assay.
Cross-service comparison. Different clocks, laboratories, sample types and normalisation pipelines produce different numbers for the same person on the same day.
Whole-body information. Acceleration in one tissue correlates only weakly with acceleration in another, and nearly all consumer tests use blood or saliva.
Any clinical conclusion. These are not diagnostic instruments, and in most jurisdictions they are not permitted to claim otherwise.
Set against that, what a defensible service can provide is a research-grade population signal, measured reliably, with uncertainty stated, that a person can watch over years. That is a real thing and a modest one.
A Reasonable Purchase Decision
If you are curious and hold the result loosely: one test from a service naming a published second or third-generation clock, using a principal-component version, in a named accredited laboratory, reporting uncertainty, returning raw data, with clear data terms. Read it as a baseline.
If you intend to track: commit to one service, one sample type, one laboratory and one annual timing, and change none of them. Consistency is worth more than picking the best option and then switching.
If the budget is limited: a standard metabolic and inflammatory panel with apoB, a cardiorespiratory fitness assessment, grip strength and honest sleep and activity tracking inform decisions more than any epigenetic test currently sold, and cost less. This is not a fashionable answer and it is the correct one for most people.
If you want the most decision-relevant version: a pace-of-ageing measure rather than a cumulative one, since rate is the part current behaviour can influence.
The market will keep producing new entrants. These seven questions will keep separating them.
The AEONNN Perspective
AEONNN does not sell a biological age test and does not require one, which makes this a genuinely disinterested evaluation. The Evidence layer does not support presenting a consumer clock result as individually actionable, so the platform reads one as context alongside metabolic, inflammatory and functional signals rather than as a trigger for a change.
Two of these seven questions carry most of the weight. Which clock and generation, since first-generation clocks were trained on chronological age and predict outcomes comparatively weakly, and whether a principal-component version was used, since standard clocks can differ by two to three years on the same sample and that exceeds any change a member could achieve in months.
The data question is the one the platform would push hardest. Methylation data carry information well beyond an age estimate, and in most jurisdictions they sit outside clinical record protections, so retention, research consent, third-party sharing and whether raw data are returned all matter more than the fee. And the platform's honest budget answer costs a service a sale: a standard panel with apoB, a fitness assessment, grip strength and sleep tracking inform more decisions for less money.
Pillar Matrix mapping
Database Matrix layers
- Quality / Formulation Layer (ConsumerLab, Labdoor)
- Evidence Layer (PubMed, Cochrane, ClinicalTrials.gov)
- Regulatory Layer (EFSA, FDA, EMA)
- Meta / Consensus Layer (JAMA, BMJ, specialty society positions)
Frequently Asked
What is the most important question to ask a biological age service?
Which clock and which generation. First-generation clocks were trained on chronological age and predict outcomes weakly; second and third generation predict outcomes considerably better.
What is a principal-component clock and why does it matter?
A version aggregating across many correlated methylation sites to reduce technical noise. Standard clocks can differ by two to three years on the same sample, which exceeds achievable change.
Is saliva as good as blood?
Generally no. Cell composition varies more in saliva and clocks calibrated on blood do not transfer cleanly. If you use saliva, stay with it rather than mixing sample types.
Should a report show uncertainty?
Yes. A single number implies a precision the method lacks. An estimate with an interval is more honest and more useful.
What should I ask about my data?
Retention period, research consent, third-party sharing, whether raw methylation data are returned, and the deletion path. This data carries information well beyond an age estimate.
Can these tests guide my decisions?
No clock is validated to indicate which change a specific person should make, and short-interval retesting largely measures the assay rather than the person.
What is a better use of the money?
A standard metabolic and inflammatory panel with apoB, a cardiorespiratory fitness assessment, grip strength and honest sleep and activity tracking. Less fashionable, more decision-relevant, cheaper.
Evidence and review
Any dosage ranges cited here reflect the ranges used in published human trials, not personal recommendations. Evidence in this field moves, so this article is reviewed quarterly and carries its last-updated date above. Nothing here is intended as medical advice, and supplementation should be discussed with a qualified clinician, particularly alongside prescribed medication or an existing condition.