WHAT THE STUDY ACTUALLY SAYS

Your smartwatch guesses your sleep stages. Against a sleep lab, the guesses drift

Wrist trackers infer light, deep and REM sleep from movement and pulse. Two validation studies find they track total sleep passably but each misjudges the stage breakdown its own way.

A consumer sleep tracker does not measure your sleep stages; it estimates them, inferring light, deep and REM sleep from wrist movement and the pulse it reads at the skin [s1]. Checked against polysomnography — the clinical gold standard that records brain waves, eye movements and muscle tone — the devices track roughly how long you sleep, but each one misjudges the breakdown of that sleep into stages in its own consistent direction [s1][s2]. The nightly hypnogram your watch shows you is the least trustworthy figure on the screen.

What the newest test found

A 2026 study in Sleep Medicine put three current wrist devices — an Apple Watch Series 7, a Fitbit Charge 5 and a Polar Vantage M2 — against home polysomnography and a research-grade actigraph in healthy adults aged 18 to 45 [s1]. Participants each wore one device: 20 with the Fitbit, 19 with the Apple Watch, 15 with the Polar [s1].

Agreement with polysomnography was, in the authors' words, generally weak, with wide limits of agreement [s1]. The errors were specific to each device and each stage. The Fitbit overestimated light sleep by an average of 57.7 minutes (P = .009) [s1]. The Apple Watch went the other way on light sleep, underestimating it by 73.8 minutes (P < .001), while overestimating REM by 30.3 minutes (P = .003) [s1]. The Polar showed a positive REM bias of 32.3 minutes (P = .026) but no significant error on the other measures [s1]. Ranked overall, the Apple Watch came out ahead, followed by the Fitbit and then the Polar [s1].

The authors' conclusion is the practical one: these devices show parameter-specific bias and limited agreement with the reference standards, and are better suited to longitudinal self-monitoring than to clinical-grade assessment of sleep architecture [s1]. In plain terms, a watch may be useful for spotting whether your sleep is trending better or worse over weeks, but the specific claim that you got, say, 42 minutes of deep sleep last night is not one the hardware can support.

The pattern is not new

This is consistent with the larger body of validation work. A 2021 study in Sleep tested seven consumer devices alongside actigraphy against polysomnography in 34 healthy young adults across three consecutive nights, one of them deliberately disrupted [s2]. It found that on an epoch-by-epoch basis the devices were good at detecting sleep — sensitivity was high across the board, at 0.93 or above [s2]. But their specificity, the ability to correctly identify the moments a person was awake, was low to medium, ranging from 0.18 to 0.54 [s2]. Comparisons of the sleep stages were, again, mixed, and the devices performed worse on the night of poorer, disrupted sleep [s2].

That combination — high sensitivity for "asleep," poor specificity for "awake" — is the signature limitation of movement-and-pulse trackers. Because people lie still for long stretches while awake in bed, a device that leans toward calling stillness "sleep" scores well on sensitivity while systematically missing wakefulness. The 2021 authors concluded that these trackers show promise for tracking sleep and wake but that their sleep-stage assessments are inconsistent [s2].

Why staging is the weak point

The gold-standard method assigns sleep stages from the electroencephalogram — the pattern of brain activity — supported by eye-movement and muscle-tone channels. A wrist device has none of those. It reconstructs stages from proxies: how much you move, and features of your heart rate and its variability derived from an optical pulse sensor. Deep sleep, light sleep and REM have characteristic heart-rate and movement signatures, but they overlap enough that any model built on those proxies will make stage-level errors, and the errors differ between manufacturers because each uses a different, proprietary algorithm [s1].

That also explains why one device over-reads REM while another under-reads light sleep on the same night: there is no shared standard for how the inference is done, and no requirement that a consumer wellness feature match a laboratory. Both studies were also small and conducted in healthy adults [s1][s2] — performance in older people, or in those with a sleep disorder, is not established by this work and could be worse.

What to watch

Whether manufacturers publish independent, stage-level validation for each new model rather than a single overall accuracy figure, since the useful question is not "does it detect sleep" but "does it get the stages right." Whether the devices are tested in the populations most likely to act on the numbers — older adults and people with insomnia or sleep apnoea — rather than healthy young volunteers. And, for a reader, whether a nightly stage score is being treated as a trend to watch over time, which the evidence supports, or as a measurement to act on, which it does not.

Sources

  1. [s1] Performance of three consumer sleep-tracking devices compared with Actigraphy and Polysomnography. Sleep Medicine, 4 June 2026. https://doi.org/10.1016/j.sleep.2026.109059
  2. [s2] Performance of seven consumer sleep-tracking devices compared with polysomnography. Sleep, 1 May 2021. https://doi.org/10.1093/sleep/zsaa291

Sources

  1. Performance of three consumer sleep-tracking devices compared with Actigraphy and Polysomnography — Sleep Medicine , June 4, 2026
  2. Performance of seven consumer sleep-tracking devices compared with polysomnography — Sleep , May 1, 2021
Related coverage