Your ring measures HRV well. The 'readiness' score built on it is another matter
Against an ECG over 536 nights, Oura's heart-rate-variability error was under 8% and Whoop's acceptable, while Garmin and Polar lagged. The recovery scores layered on top stay proprietary and largely unvalidated.
| Group | Value (%) |
|---|---|
| Oura Gen 4 | 5.96 |
| Oura Gen 3 | 7.15 |
| Whoop 4.0 | 8.17 |
| Garmin Fenix 6 | 10.52 |
| Polar Grit X Pro | 16.32 |
Heart-rate variability — the small beat-to-beat variation in the timing of the heartbeat — is the signal underneath most of the "readiness," "recovery" and "body battery" scores that consumer wearables now show each morning. The good news from recent validation work is that a decent device measures HRV reasonably well overnight. The caveat is that the branded daily score sitting on top of it is a proprietary calculation, and the evidence that it means what it claims is far thinner than the evidence for the raw measurement [s1][s3].
The raw measurement: device-dependent, but often good
A validation study worn simultaneously with an electrocardiogram across 536 nights in 13 adults compared five wearables [s1]. For nocturnal HRV, the two Oura rings were most accurate — Oura Gen 4 had a mean absolute percentage error of 5.96% (concordance 0.99) and Gen 3 an error of 7.15% (concordance 0.97) [s1]. Whoop 4.0 was acceptable at 8.17% error (concordance 0.94), while the Garmin Fenix 6 (10.52%) and Polar Grit X Pro (16.32%) showed poorer agreement [s1].
Resting heart rate was measured even more tightly: Oura Gen 3 and Gen 4 landed within about 2% of the ECG, Whoop within 3%, and Polar somewhat worse; Garmin was excluded from the resting-heart-rate analysis for methodological inconsistencies [s1]. A broader review of smart rings similarly put HRV agreement very high (r² = 0.980) across the studies it pooled [s3]. So the input is real: on a good device, overnight HRV and resting heart rate are genuine measurements, not decorations.
Does HRV track how you feel? Modestly
If the number is accurate, the next question is whether it tells you anything useful day to day. A 14-day study of 41 adults taking standardised morning HRV readings found that higher HRV was associated with better self-reported sleep (β = 0.510), lower fatigue (β = 0.281) and reduced stress (β = 0.353), even after adjustment [s2]. It found no association with muscle soreness [s2].
The authors were careful about how much to make of this: the effect sizes were modest and individual variability was substantial, so HRV is best read as one contextual signal tracked over time in a given person, not a precise daily verdict on recovery [s2]. That is a meaningful distinction from how the scores are marketed, and it fits the more ambivalent user experience described in a qualitative study of people who felt their tracker had become "a toxic relationship".
The score itself is a black box
Here is the gap. "Readiness" (Oura), "recovery" (Whoop) and "Body Battery" (Garmin) are composite scores that blend HRV, resting heart rate, sleep and other inputs through an algorithm each company keeps proprietary. The systematic review of smart rings found that 89% of the studies it examined relied on such proprietary algorithms, which cannot be independently inspected or reproduced [s3].
That opacity has two consequences. First, a validated input does not guarantee a validated output: a score can be built on accurate HRV and still combine it in a way no external study has tested against any outcome. Second, when Garmin's HRV agreement is among the weakest of the devices tested [s1], a Body Battery reading resting partly on that signal inherits the uncertainty — while presenting a single, confident-looking number. The published evidence validates the sensors far better than it validates the scores.
HRV is noisier than heart rate for a reason
It is not an accident that every device measured resting heart rate more tightly than heart-rate variability [s1]. Resting heart rate is a single averaged number; HRV is a statistic about the tiny differences between consecutive beats, so it amplifies any small timing error the optical sensor makes, and it is far more sensitive to when and how it is measured — posture, breathing, a slightly restless stretch of sleep. That is why HRV also swings more from day to day within the same healthy person, as the 14-day study's substantial individual variability showed [s2]. A single night's HRV is therefore weak evidence of anything; a baseline built over weeks, and a sustained departure from it, is where the metric earns its keep [s1][s2].
This also cautions against comparing HRV between people or between devices. Absolute HRV varies enormously by age and physiology, and the device-to-device error the validation exposed — from under 6% on the best ring to over 16% on the weakest device — means two people's numbers, or the same person's numbers on two brands, are not on a common scale [s1]. The only fair comparison is a person against their own trend on one device.
What it means for a user
Treat the raw metrics as the trustworthy part, and even then device-by-device: overnight resting heart rate and HRV on a well-validated device are real, and a downward drift in HRV sustained over days can be a reasonable prompt to rest [s1][s2]. Treat the branded daily score as a convenience summary, not a validated recovery verdict — its formula is undisclosed, its output untested against hard outcomes, and its precision oversold relative to the modest, individual-specific way HRV actually relates to how you feel [s2][s3]. Watch the trend in the underlying number, not the single digit the app puts in front of it.
Sources
- Validation of nocturnal resting heart rate and heart rate variability in consumer wearables — Physiological Reports , August 1, 2025
- Associations Between Daily Heart Rate Variability and Self-Reported Wellness: A 14-Day Observational Study in Healthy Adults — Sensors , July 15, 2025
- Smart Ring in Clinical Medicine: A Systematic Review — Biomimetics , December 5, 2025
Temperature-sensing wearables can flag ovulation — within a few days, not to the day
A network meta-analysis put pooled accuracy for the fertile window at 0.88, best in the three days around ovulation. A wrist-temperature study hit ovulation within three days about 78% of the time.
Your watch's VO2max is a decent estimate for ordinary fitness, shakier at the elite end
A review found wearables gave valid maximal-oxygen-uptake estimates in most studies of untrained and recreational exercisers. A separate meta-analysis pooled a 0.83 correlation, with high variability.
What a smart scale's body-fat number is worth against the reference scan
A smartwatch bioimpedance sensor tracked body-fat percentage against DXA with about 14% average error — but muscle-mass agreement was weak, and even sitting versus standing shifted the reading.
Wrist heart-rate sensors are getting better, but they still disagree during exercise
Head-to-head against an ECG chest strap, popular optical wrist devices differed substantially, and one graded test found a smartwatch running a couple of beats low with a 28-bpm spread in its limits of agreement.