Skip to main content
Metric JournalRecovery

How Apple Watch Recovery Scores Work

From sensor samples to a 0–100 number: the full calculation chain and where uncertainty enters.

Share on XShare on Threads

Apple Watch recovery scores work by reading authorized HealthKit records, comparing selected signals with a recent personal baseline, weighting the deviations, smoothing noise, and mapping the result to a simple scale. Apple supplies measurements and separate features such as Vitals, Sleep Score, and Training Load; third-party apps decide which of those inputs become a recovery score.

Editorial note

Published 2026-07-16. Last reviewed 2026-07-16. Evidence, interpretation, and Metric’s implementation are identified separately below.

Metric workout summary showing the measurements behind training context
A recovery score is the last step in a pipeline; the recorded workouts and trends behind it remain essential context. Screen shown with representative data.

The six-stage calculation chain

  1. Permission and collection: the app requests specific HealthKit categories and receives only the records available to it.
  2. Quality control: the app filters values by date, source, plausibility, sample count, and freshness.
  3. Baseline construction: it summarizes a recent period for that person, often with rolling averages or robust statistics.
  4. Deviation scoring: it determines whether today differs from baseline and in which direction.
  5. Combination: it weights cardiac, sleep, workload, and sometimes subjective inputs.
  6. Presentation: it smooths or caps the result, maps it to a scale, and assigns a label such as recovered, balanced, or strained.

Where two apps begin to disagree

  • Input selection: one app may emphasize overnight HRV while another emphasizes workload.
  • Measurement choice: Apple Health stores SDNN for HRV, while another app may calculate RMSSD from its own recording.
  • Baseline length: a seven-day comparison reacts faster than a 60- or 90-day window.
  • Weighting: a short night can dominate one formula and barely move another.
  • Missing data: one app may withhold the score, another may reweight the remaining inputs, and another may carry a prior value forward.
  • Score direction: in some apps a higher number means more recovered; in a strain view a higher raw load may mean the opposite.
  • Intended decision: a score designed for endurance training is not necessarily designed for illness detection or general wellbeing.

Apple’s own building blocks

Apple’s Vitals feature establishes typical ranges after sufficient overnight wear and flags multiple outliers. Sleep Score uses duration, bedtime consistency, and interruptions. Training Load compares seven days of workout intensity and duration with the prior 28 days. These are transparent examples of different questions producing different summaries; none alone is a universal recovery value.

What the evidence says

A 12-week study tracked 43 endurance athletes using training, nutrition, sleep, HRV, and subjective wellbeing. Group models improved prediction of perceived morning recovery and HRV change over a simple baseline, but person-level errors varied substantially. That is exactly why a product should show trend, confidence, and contributors instead of presenting a score as ground truth.

How Metric’s score works

Metric’s current readiness/strain calculation starts from authorized HealthKit workouts. It converts workouts into daily load, maintains a faster-changing acute load and a slower-changing chronic load, and gives some weight to load accumulated today. Rest days allow short-term load to decay faster. Those components are passed through bounded curves and combined into a score from 0 to 100, where a higher number means lower modeled strain.

This design answers a workload question: how much recent training demand are you carrying relative to your longer pattern? Sleep, HRV, symptoms, and soreness can still change the sensible decision. Metric therefore presents the result as a readiness/strain guide and not as a direct reading of recovery tissue, immune status, or injury risk.

How to audit any recovery score

  • Can you name the major inputs and their measurement windows?
  • Can you see when an input is missing or stale?
  • Does the app explain whether higher means better, more load, or more risk?
  • Can you return to the underlying trend rather than seeing only the score?
  • Does the score stabilize after enough baseline data, or jump because of single samples?
  • Does the recommendation account for the activity you actually plan to do?

Limitations

  • HealthKit permission gaps and device non-wear can make the input set incomplete.
  • A personal baseline can be stable but still reflect an unusual training period, illness, or travel.
  • Smoothing improves stability but can delay the response to a real change.
  • A 0–100 scale suggests precision that the sensors and model may not support at the single-point level.
  • Recovery is multidimensional; no wrist-derived score observes every relevant system.

Sources and further reading

Continue exploring

Health information disclaimer

Metric is a wellness product. This article is educational and does not provide medical advice, diagnosis, or treatment. Wearable measurements and app-generated scores are estimates; discuss symptoms, unusual readings, medication effects, and changes to a care or training plan with an appropriate healthcare professional.

See workout load and recovery context in Metric
← Back to Metric Journal