AEONNN How It Works Pillars Membership FAQ Journal AEONNNian Access Request Early Access

Wearable Data and Longevity: Turning Signals Into Action

Which wearable metrics are validated, which are proprietary estimates, and the four that are actually worth acting on. Plus the failure mode where tracking makes the thing it measures worse.

7 min read

The Short Answer

Wearables generate a great deal of data and a small amount of information. Some metrics are well validated against reference methods, some are proprietary composites whose derivation is undisclosed, and a few are effectively invented. The useful discipline is knowing which is which, then acting on the four or five that survive. The failure mode worth naming early is that tracking can degrade what it measures: sleep anxiety driven by sleep scores is documented well enough to have a name.

Metric by Metric

MetricValidationWorth acting on
Steps and activity volumeGoodYes
Resting heart rateGoodYes, as a trend
Heart rate during steady exerciseGood for chest strap, moderate for wrist opticalYes
Heart rate variabilityReasonable for the measurement, highly variable in interpretationAs a multi-week trend only
Sleep durationModerate; better than self-reportYes
Sleep stagingPoor to moderate against polysomnographyRarely
Sleep regularityGood, and derived from timing rather than stagingYes, strongly
Blood oxygen saturationVariable; affected by skin tone, motion and fitOnly for a persistent pattern
Skin temperature trendReasonable for relative changeSometimes, for illness and cycle tracking
Estimated VO2 maxModerate; useful for trend, not absolute valueAs a trend
Readiness and recovery scoresProprietary composites, undisclosed derivation, not independently validatedLoosely at best
Stress scoresUsually a heart rate variability derivative with an interpretive layerNo

The proprietary composites are where most user attention goes and where the least validation exists. A readiness score combines several inputs by an undisclosed formula and produces a single number, and neither its derivation nor its predictive validity is published. It may correlate with something useful. Nobody outside the company can say.

The Four Worth Acting On

Sleep regularity. The strongest signal any consumer wearable produces, and the most neglected. Consistency of sleep and wake timing predicts mortality and cardiometabolic outcomes in large cohort analyses, in some analyses more strongly than duration. It is derived from timing rather than from staging, which is why it is reliable. Acting on it means fixing wake time first, since it anchors everything downstream.

Resting heart rate trend. A slow decline with training and a sustained rise with illness, poor sleep, alcohol or accumulating load. The trend is informative and single days are not.

Activity volume, especially low intensity. Total daily movement associates strongly with mortality, with most of the benefit accruing at the low end of the range. Moving from very low to moderate activity carries more benefit than moving from moderate to high.

Estimated cardiorespiratory fitness trend. The absolute number is imprecise and the direction over quarters is among the most meaningful things a person can track, given how strongly fitness predicts mortality.

Notably, all four are trends rather than daily values, and none is a proprietary score.

Heart Rate Variability, Carefully

HRV deserves separate handling because it is simultaneously the most discussed and most misused wearable metric.

What it measures is real: beat-to-beat variation reflecting autonomic balance, with higher values generally indicating greater parasympathetic influence. Population associations with cardiovascular outcomes exist.

The problems are all in interpretation. Absolute values vary several-fold between healthy individuals, so comparison to anyone else is meaningless. Day-to-day variation within a person is large, driven by alcohol, late meals, illness, position, breathing pattern, hydration and measurement timing. Measurement method matters, and values from different devices are not comparable. And higher is not always better: HRV can be elevated in overreached athletes and in some clinical states.

What HRV supports is a personal multi-week trend, read alongside training load, sleep and alcohol. A sustained decline over several weeks alongside worse sleep and rising resting heart rate is a real signal about accumulating load. A single low morning means very little, and reading it as a verdict on the day ahead is the most common misuse.

The Feedback Loop Problem

Tracking is not neutral, and this is the least discussed limitation of the whole category.

Sleep score anxiety. Sleep is unusually susceptible to attention: worrying about it degrades it. A poor score in the morning that shapes expectations for the day, or anxiety at bedtime about what tonight's score will be, produces exactly the outcome it measures. Clinicians working in sleep medicine have begun encountering this specifically, and the remedy is sometimes to stop tracking.

Score chasing. Optimising a proprietary composite whose derivation is undisclosed means optimising an unknown function, which may or may not correspond to health.

Displaced attention. Time spent reviewing dashboards is not time spent training, cooking or sleeping. The measurement can crowd out the thing measured.

False reassurance. Good scores can substitute for actual assessment. A wearable does not measure blood pressure, lipids, glucose regulation or any of the things that most predict cardiovascular outcomes.

A reasonable discipline: review weekly rather than daily, look at trends rather than values, ignore proprietary composites, and stop tracking any metric that is making the underlying behaviour worse.

What Wearables Cannot Tell You

The gap between wearable data and what determines long-term outcomes is wide and worth stating plainly.

No wearable measures apolipoprotein B, lipoprotein(a), HbA1c, inflammatory markers, kidney or liver function, thyroid status or nutrient status, and those are the measurements that most inform long-term risk. Continuous glucose monitoring is the partial exception and is a separate class of device with its own interpretive difficulties in non-diabetic use.

Wearables also do not measure strength, which is among the strongest predictors of function in later life. A grip dynamometer costs very little and predicts more than most of a wearable's output.

The honest framing is that wearables are good at behaviour and trends, particularly sleep timing and activity, and blind to biochemistry. A person tracking only wearable data has a detailed picture of their habits and none of their physiology.

Using the Data Sensibly

A defensible practice, given all of the above:

Pick four metrics and ignore the rest. Sleep regularity, resting heart rate trend, activity volume and a fitness estimate. Everything else is noise or entertainment.

Review weekly. Daily review invites reacting to variation that carries no information.

Pair with periodic biochemistry. Wearables handle behaviour; a blood panel handles physiology. Neither substitutes for the other, and a person doing only one of the two has half a picture.

Use it to detect change, not to grade yourself. The value is in noticing a multi-week drift you would otherwise have missed, which is a different use from receiving a daily verdict.

Stop if it is making things worse. This applies most to sleep, where the tracking can become the problem. Recognising that is a reasonable outcome rather than a failure.

Used this way, a wearable is genuinely useful and it is a modest instrument rather than a health platform. The metrics that survive scrutiny are few, unglamorous and trend-based, which is a fair description of most things that work in this field.

The AEONNN Perspective

The Real-Time User layer is where wearable data enters AEONNN, and the platform's discipline is to use the validated signals and disregard the proprietary composites. Sleep regularity, resting heart rate trend, activity volume and fitness estimates are what the Contingency layer reads when it adapts a stack to travel, illness or accumulating load.

Heart rate variability is used as a personal multi-week trend and never as a daily verdict, because between-person variation is several-fold and day-to-day variation within a person is large. That is a Quality layer judgement about the measurement rather than a claim about the physiology.

It maps across Pillar 10, Pillar 9 and Pillar 4. The limitation the platform states plainly is that wearables are blind to biochemistry: no device measures apolipoprotein B, HbA1c, inflammatory or nutrient status, and those inform long-term trajectory more than anything on a wrist. A member tracking only wearable data has a detailed record of behaviour and none of physiology.

Database Matrix layers

  • Real-Time User Layer (wearable and adherence signals)
  • Evidence Layer (PubMed, Cochrane, ClinicalTrials.gov)
  • Quality / Formulation Layer (ConsumerLab, Labdoor)
  • Population Layer (UK Biobank, NHANES)

Frequently Asked

Which wearable metrics are actually reliable?

Steps and activity volume, resting heart rate, sleep duration and especially sleep regularity, plus fitness estimates read as trends. Sleep staging is poor against polysomnography and proprietary readiness scores are unvalidated.

What is the most useful wearable metric?

Sleep regularity. Consistency of sleep and wake timing predicts mortality and cardiometabolic outcomes in large cohorts, in some analyses more strongly than duration, and it is derived from timing rather than staging.

How should I use heart rate variability?

As a personal multi-week trend read alongside training load, sleep and alcohol. Absolute values vary several-fold between people, day-to-day variation is large, and a single low morning means very little.

Are readiness and recovery scores meaningful?

They are proprietary composites with undisclosed derivation and no independent validation. They may correlate with something useful, and nobody outside the company can say what.

Can tracking make sleep worse?

Yes. Sleep is susceptible to attention, and anxiety about scores can produce the poor sleep it measures. Sleep clinicians encounter this specifically, and stopping tracking is sometimes the remedy.

What can wearables not measure?

Apolipoprotein B, lipoprotein(a), HbA1c, inflammatory markers, kidney and liver function, thyroid and nutrient status, and strength. Those inform long-term risk more than anything a wrist device reports.

How often should I review the data?

Weekly, looking at trends. Daily review invites reacting to variation that carries no information.

Evidence and review

Any dosage ranges cited here reflect the ranges used in published human trials, not personal recommendations. Evidence in this field moves, so this article is reviewed quarterly and carries its last-updated date above. Nothing here is intended as medical advice, and supplementation should be discussed with a qualified clinician, particularly alongside prescribed medication or an existing condition.

Continue Reading

Membership

Reading about longevity and Biological Age is not the same as knowing where you stand.

AEONNN organizes an article like this one against your own profile. Origin works through Discovered Mode, building your Pillar Matrix from the context you provide. Evolution adds Synched Mode, so supported wearable, Apple Health and laboratory data inform the same reasoning.

AEONNN turns knowledge like this into a protocol that is yours.

Private Early Access opens in August. Public launch follows in September.

By requesting access, you agree to receive AEONNN launch and membership communications. You may unsubscribe at any time. Privacy Policy · Consumer Health Data Privacy Notice

Back to the Journal →