Wearables flood into research. The 'digital biomarker' is still mostly a promise
An umbrella review of 42 reviews found only two wearable measures — heart rate and atrial fibrillation — with comparably strong support. Getting any device into a trial is harder than buying it.
"Digital biomarker" has become one of the most reached-for phrases in health technology: the idea that a signal from a wearable can stand in for a clinical measurement, track a disease, or serve as an endpoint in a drug trial. The 2026 evidence suggests the category is real but early — a handful of measures are genuinely well-supported, many more are promising but unproven, and the practical work of getting any wearable into a rigorous study is much harder than the marketing implies [s1][s2][s3].
How little is strongly validated
The most sobering number comes from an umbrella review in PLOS Digital Health, which synthesised 42 existing reviews spanning more than 30 brands and over 150 device series, assessing wearables for monitoring people with long COVID and related post-acute infection syndromes [s3]. It found highly variable review quality, substantial heterogeneity in device performance across biometrics and populations, and limited evidence for clinical utility [s3]. Of the 17 biometrics it evaluated, only two — heart-rate measurement and atrial-fibrillation detection — had comparably stronger support [s3].
That is a narrow base. It does not mean other signals are worthless, but it does mean the confident clinical claims attached to sleep stages, stress, recovery, respiratory rate and the rest are running ahead of the validation. A review guiding clinicians through consumer wearables makes the same point structurally: accuracy and validity vary, and any use of wearable data needs practical caveats about the device, the setting and the population rather than blanket trust [s2]. It is the gap between a measurement that exists and a measurement that has earned a clinical role — the same gap we saw in a wrist-sensor aging clock built on photoplethysmography.
Why getting a wearable into a trial is hard
Even where a signal is valid, operationalising it is not simple. An experience-based tutorial from investigators who run large wearable trials lays out why: incorporating consumer devices "is not as straightforward as simply purchasing devices and providing them to participants" [s1]. It requires navigating a complex engineering ecosystem, with decisions about device selection, data-management software, and participant engagement and technical support [s1].
Each of those is a place a study can quietly fail. Manufacturers change firmware and restrict data access mid-study; different form factors — watch, band, ring — expose different data at different resolutions; and a metric that looks like a clean number in an app may arrive as irregular, proprietary, or aggregated data that an investigator cannot fully interrogate [s1]. The tutorial's contribution is a structured framework for matching device and data tools to a study's objectives [s1], which is itself a sign that the field is still assembling its basic methods.
The reproducibility problem underneath
These operational hurdles compound a scientific one. Because so many wearable metrics are generated by proprietary algorithms that change over time and cannot be inspected, a "digital biomarker" measured on one device in one year may not be the same quantity measured on another device later. That is part of why claims are so hard to compare across studies — the problem a reporting checklist tried to address for the wearable-coaching models we covered in a tuned LLM's performance on sleep and fitness. Heterogeneity and weak external validation recur as the limiting factors across these syntheses [s2][s3].
The measurement is not the endpoint
There is a deeper conceptual gap worth naming. A digital biomarker has to clear two separate bars: that the device measures the signal accurately, and that the signal means something clinically — that it tracks a disease, predicts an outcome, or responds to a treatment in a way that matters. The wearable literature has made real progress on the first bar and much less on the second. The umbrella review's finding that only heart rate and atrial fibrillation had comparably strong support is a finding about clinical utility, not merely accuracy [s3], and the cardiovascular guide's insistence on practical caveats is a warning that an accurate number can still be the wrong number to act on [s2].
Regulators have drawn the same line from the other direction: a wearable can measure a physiological parameter without being allowed to make a claim about what it means, the distinction at the heart of the FDA's general-wellness guidance. For a drug trial, that second bar is the whole game — an endpoint that regulators will accept has to be shown to capture something patients feel or clinicians can change, which is a far higher standard than a good correlation with a reference sensor.
What it means
For the field, the realistic reading is that wearables are becoming useful research instruments for a few well-validated signals and for decentralised data collection, while the broader "digital biomarker" vision remains aspirational and unevenly evidenced [s1][s2][s3]. For a reader encountering a headline that a wearable "predicts" or "detects" a disease, the useful questions are the ones these reviews kept returning to: which specific biometric, validated against what reference, in which population, on which device — and whether anyone has reproduced it. Heart rate and atrial fibrillation can currently answer those questions; most of the rest cannot yet [s3].
Sources
- Device and Data Access Considerations for Digital Health Studies Using Consumer-Grade Wearables: An Experience-Based Tutorial — Clinical and Translational Science , August 1, 2026
- A guide to consumer-grade wearables in cardiovascular clinical care and population health for non-experts — npj Cardiovascular Health , September 2, 2025
- Can consumer wearables support outpatient health monitoring for patients with post-acute infection syndromes? A systematic umbrella review of accuracy, validity, and clinical utility data — PLOS Digital Health , June 8, 2026
Can a wearable catch illness before you feel it? The evidence is intriguing and shaky
Smart-ring studies report flagging COVID nearly three days before symptoms and IBD flares weeks ahead. But two-thirds of that research carries a high risk of bias, and clinical usefulness is largely unproven.
Smart toilets and passive vitals: promising hardware, thin independent evidence
Toilet seats and mirrors that read your heart and blood pressure are real engineering. The published validation is mostly small, single-lab or company-run — and outcome evidence is absent.
What the Apple Watch ECG can and cannot detect, and the clearance it rests on
The FDA cleared it to sort one rhythm from one other. It cannot detect a heart attack, and about one reading in eight comes back inconclusive.
A tuned LLM outscored experts on sleep and fitness exams. Then came the hard part
Google's PH-LLM beat sampled human experts on multiple-choice tests but only matched them on real cases. A new reporting checklist published the same month explains why such claims are hard to compare.