A tuned LLM outscored experts on sleep and fitness exams. Then came the hard part
Google's PH-LLM beat sampled human experts on multiple-choice tests but only matched them on real cases. A new reporting checklist published the same month explains why such claims are hard to compare.
Two papers published in August approach the same problem from opposite ends. One reports a large language model fine-tuned to interpret wearable data and give sleep and fitness advice, and reports it doing well [s1]. The other is a consensus reporting checklist for exactly this kind of study, built because the field's claims have become impossible to compare with each other [s2].
Read together they say something more useful than either says alone.
What the model is
The Personal Health Large Language Model, or PH-LLM, is a version of Google's Gemini fine-tuned for text understanding and reasoning over aggregated daily-resolution numerical sensor data — the kind of data a consumer wearable produces [s1]. The stated gap it targets is real: LLMs have been evaluated extensively on clinical text, and much less on personalised health monitoring from device data [s1].
The researchers built three benchmark datasets to test three different things: expert domain knowledge, generation of personalised insights and recommendations, and prediction of self-reported sleep quality from longitudinal data [s1].
What it scored
On multiple-choice examinations, PH-LLM exceeded a sample of human experts — 79% versus 76% in sleep medicine, and 88% versus 71% in fitness [s1].
Those are the numbers most likely to be quoted, and they are the least informative ones. Multiple-choice performance measures retrieval of codified knowledge, which is the task LLMs are structurally best at. The three-point gap in sleep medicine is narrow enough that the composition of the human comparison sample matters a great deal.
The harder evaluation involved 857 real-world case studies. There, PH-LLM performed similarly to human experts on fitness-related tasks and improved over the base Gemini model in providing personalised sleep insights [s1]. That is a more modest claim, and a more meaningful one: parity with experts on open-ended cases rather than superiority on exams.
Finally, the model predicted self-reported sleep quality using a multimodal encoding of wearable sensor data, which the authors present as evidence it can contextualise sensor modalities rather than merely describe them [s1].
The paper releases its datasets, rubrics and benchmark performance to let other groups build on the work [s1] — a genuine contribution, and one that the second August paper suggests is the exception rather than the norm.
Why the second paper exists
The CHART Statement — Chatbot Assessment Reporting Tool — was published in JAMA Network Open on 1 August [s2]. Its stated rationale is blunt: the rise in chatbot health advice studies has been accompanied by heterogeneity in reporting standards, which affects how interpretable those studies are [s2].
CHART was developed through a comprehensive systematic review that identified variation in how such studies are conducted, reported and designed; a draft checklist was then revised through an international modified asynchronous Delphi consensus process involving 531 stakeholders, three synchronous panel consensus meetings of 48 stakeholders, and pilot testing [s2]. The result is 12 items and 39 subitems [s2].
The items are worth listing in outline, because they map the things that routinely go unreported: model identifiers, model details, prompt engineering, query strategy, performance evaluation, sample size, data analysis, disclosures, funding, ethics, protocol and data availability [s2].
The gap the checklist exposes
Consider what "model identifiers" and "prompt engineering" mean in practice. A study reporting that a chatbot answered health questions at some accuracy rate is describing a system consisting of a base model, a specific version of that model, a system prompt, a query format and a decoding configuration. Change any of those and the number changes. Commercial models are also updated without notice, so a result obtained in March may not reproduce in September on what is nominally the same product.
CHART's authors position the checklist as support for clinicians, researchers, editors, peer reviewers and readers in reporting, understanding and interpreting these findings [s2]. The word "readers" is doing real work there. Chatbot health advice studies get press coverage far out of proportion to their methodological maturity, and the headline number — an accuracy percentage, an exam score — is almost always the least transferable part of the result.
What neither paper claims
PH-LLM was evaluated on benchmarks and case studies. Neither the paper nor this account describes a trial in which people used it and their sleep or fitness measurably changed. The distinction between "the model produces advice experts rate as good" and "the advice makes people healthier" is the distinction that matters for anyone deciding whether a wearable's AI coach is worth attending to, and it remains unaddressed.
The CHART statement, for its part, is a reporting guideline. It standardises how studies are described; it does not make the underlying studies better designed, and it cannot make a chatbot's advice safe. Reporting checklists in other fields — CONSORT for trials, PRISMA for reviews — took years to change practice and never eliminated poor studies. They made poor studies easier to recognise.
What to watch
Whether journals adopt CHART as a submission requirement is the concrete near-term question, because that is the mechanism by which reporting guidelines actually bite. The second is whether the consumer-facing versions of models like PH-LLM are evaluated in real use rather than on benchmarks — and whether those evaluations, when they come, report the details CHART asks for.
Sources
- A personal health large language model for sleep and fitness coaching — Nature Medicine, 2025-08-14
- Reporting Guideline for Chatbot Health Advice Studies: The CHART Statement — JAMA Network Open, 2025-08-01
Sources
- A personal health large language model for sleep and fitness coaching — Nature Medicine , August 14, 2025
- Reporting Guideline for Chatbot Health Advice Studies: The CHART Statement — JAMA Network Open , August 1, 2025
More on
NICE reviewed 30 digital mental health tools. The same evidence gaps kept recurring
A cross-sectional analysis of NICE evaluations found 78 supporting studies behind 30 technologies — and consistent holes in comparators, cost of delivery and adverse-event reporting.
A small insomnia trial asked when in the day a sleeping pill helps — and when it hurts
Standard questionnaires found no daytime difference between suvorexant and placebo. Four-times-daily smartphone prompts found fatigue worse in the morning and better later. The trial had 40 people.
Tuning daytime and evening light did not improve older adults' sleep
In a 12-person crossover trial, blue-enriched daytime light with reduced evening blue did not cut time spent awake at night, though it nudged a brainwave marker of sleep.
Light alone barely moved young adults' clocks; a combined package did
Two small pilot studies in 18-to-25-year-olds found morning bright light on its own shifted circadian timing little, while adding a timed schedule and evening blue-blockers advanced it about 35 minutes.