A smartwatch matched the sleep lab on severe apnea. The threshold is the catch.
In 152 adults tested against two nights of polysomnography, the Galaxy Watch reached an AUROC of 0.94 for moderate-to-severe sleep apnea — with 66.7% specificity at the setting it ships with.
| Group | Value (%) |
|---|---|
| Default threshold 15, sensitivity | 94.1 (84.1 to 98) |
| Default threshold 15, specificity | 66.7 (51 to 79.4) |
| Optimised threshold 25.95, sensitivity | 82.4 (69.7 to 90.4) |
| Optimised threshold 25.95, specificity | 94.9 (83.1 to 98.6) |
Obstructive sleep apnea remains substantially underdiagnosed because the definitive test — in-laboratory polysomnography — is resource-intensive and not readily accessible to everyone who needs it [s1]. That gap is the commercial opening consumer wearables have been aiming at for years, and the evidence supporting them has generally been thinner than the marketing.
A prospective validation study published in the Journal of Clinical Sleep Medicine on 27 August is a genuine addition to that evidence, and it is worth reading closely for what it establishes and what it leaves open [s1].
The test
Researchers enrolled 152 adults aged 22 or older who either had a prior diagnosis of moderate-to-severe obstructive sleep apnea or a high pre-test likelihood of it, defined as a STOP-Bang score of 3 or more [s1]. Participants wore a Samsung Galaxy Watch through an initial in-lab polysomnography night, then at least three watch-only nights at home, then a second in-lab night; 147 completed both laboratory nights, and the study accumulated 1,850 hours of in-lab polysomnography sleep [s1].
Two reference standards were used. The first was the familiar apnea-hypopnea index, with moderate-to-severe disease defined as an AHI of 15 or more events per hour [s1]. The second was hypoxic burden, which integrates the depth, duration and frequency of oxygen desaturations and which the authors note has been shown to reflect cardiometabolic risk and disease severity better than AHI does [s1].
What it found
Against the AHI standard, the watch achieved an area under the receiver operating characteristic curve of 0.94 (95% CI, 0.889–0.980) for detecting moderate-to-severe sleep apnea [s1]. That is a strong discrimination figure by any measure.
The threshold choice is where it gets interesting. At the device's default estimated-AHI cut-off of 15, sensitivity was 94.1% (95% CI, 84.1–98%) and specificity 66.7% (95% CI, 51–79.4%) [s1]. At a cohort-optimised estimated-AHI threshold of 25.95, the trade-off reversed: sensitivity 82.4% (95% CI, 69.7–90.4%), specificity 94.9% (95% CI, 83.1–98.6%) [s1].
Within the group defined as high-risk by hypoxic burden, the watch reached 100% sensitivity and 100% specificity at the default threshold, which the authors describe as robust discrimination between hypoxic-burden-defined low- and high-risk categories [s1].
Reading the specificity number
A default setting that catches 94.1% of moderate-to-severe cases and wrongly flags roughly a third of those without it is a defensible design for a screening device: in screening, a false alarm costs a clinic visit, a missed case costs years of untreated disease. But that trade is only defensible if the false alarms land somewhere that can absorb them.
Two limits sit on top of the specificity figure. The confidence interval around it is wide — 51% to 79.4% [s1] — because the number of true negatives in a cohort selected for high pre-test probability is small by construction. And that selection is itself the study's largest constraint. Everyone enrolled either had diagnosed moderate-to-severe apnea already or scored 3 or more on STOP-Bang [s1]. Performance in a general population, where prevalence is far lower and positive predictive value falls accordingly, is not what was measured here.
The perfect discrimination within the hypoxic-burden high-risk group is the most striking result in the paper and also the one most in need of replication. It rests on a subgroup of a cohort of 152.
The wider framing
Sleep published a Perspective on 28 August that names the problem this kind of result runs into on its way to the public. Sleep, it argues, has moved from an overlooked biological process to a widely recognised component of health, and that recognition has created a translational challenge: research increasingly treats sleep as a multidimensional, context-dependent construct, but those nuances get attenuated as findings travel to clinicians, policymakers, commercial stakeholders and the public [s2]. It argues for communication that preserves evidentiary nuance and acknowledges uncertainty, so that sleep's growing prominence matures rather than becoming "simply the latest target of health optimization and overstatement" [s2].
The distance between "AUROC 0.94" and "your watch can diagnose sleep apnea" is exactly that attenuation. The study's own conclusion is narrower than either: the findings support consumer wearables as a scalable, accessible screening tool to improve risk identification, triage and referral for definitive diagnostic evaluation [s1]. Referral is the operative word. Nothing here replaces a diagnostic study; what it supports is a device that tells more people they should get one.
The trial was registered on ClinicalTrials.gov (NCT06603441, first posted 19 September 2024) [s1].
Sources
- [s1] "Smartwatch-based detection of moderate-to-severe and high-risk obstructive sleep apnea," Journal of Clinical Sleep Medicine, 27 August 2026. https://doi.org/10.1007/s44470-026-00159-8
- [s2] "Sleep in Context," Sleep, 28 August 2026. https://doi.org/10.1093/sleep/zsag236
Sources
- Smartwatch-based detection of moderate-to-severe and high-risk obstructive sleep apnea — Journal of Clinical Sleep Medicine , August 27, 2026
- Sleep in Context — Sleep , August 28, 2026
More on
A quarter of Chileans screen positive for sleep apnoea risk — on questionnaires alone
Only three of the 15 studies pooled in a Chilean meta-analysis used sleep testing at all. A Lima clinic analysis shows why that shortcut matters.
Sleep questionnaires are good at ruling out, and poor at ruling in
The best-validated sleep apnea screener works by being wrong in a specific direction. A 47-study meta-analysis of 26,547 people shows what a questionnaire can and cannot conclude about you.
The apnea number surgeons use may be the wrong one for predicting complications
In 2,286 patients with sleep apnea undergoing major surgery, a measure of how deeply and how long oxygen fell — not the apnea-hypopnea index — tracked 30-day cardiovascular events and death.
Sleep apnea plus insomnia raised hypertension risk — but only in one insomnia subtype
In a Pennsylvania cohort followed 7.5 years, the combination carried the highest risk when insomnia came with objectively short sleep. Insomnia with normal sleep duration showed no significant association.