ANALYSIS

A tuned LLM outscored experts on sleep and fitness exams. Then came the hard part

Google's PH-LLM beat sampled human experts on multiple-choice tests but only matched them on real cases. A new reporting checklist published the same month explains why such claims are hard to compare.

Two papers published in August approach the same problem from opposite ends. One reports a large language model fine-tuned to interpret wearable data and give sleep and fitness advice, and reports it doing well [s1]. The other is a consensus reporting checklist for exactly this kind of study, built because the field's claims have become impossible to compare with each other [s2].

Read together they say something more useful than either says alone.

What the model is

The Personal Health Large Language Model, or PH-LLM, is a version of Google's Gemini fine-tuned for text understanding and reasoning over aggregated daily-resolution numerical sensor data — the kind of data a consumer wearable produces [s1]. The stated gap it targets is real: LLMs have been evaluated extensively on clinical text, and much less on personalised health monitoring from device data [s1].

The researchers built three benchmark datasets to test three different things: expert domain knowledge, generation of personalised insights and recommendations, and prediction of self-reported sleep quality from longitudinal data [s1].

What it scored

On multiple-choice examinations, PH-LLM exceeded a sample of human experts — 79% versus 76% in sleep medicine, and 88% versus 71% in fitness [s1].

Those are the numbers most likely to be quoted, and they are the least informative ones. Multiple-choice performance measures retrieval of codified knowledge, which is the task LLMs are structurally best at. The three-point gap in sleep medicine is narrow enough that the composition of the human comparison sample matters a great deal.

The harder evaluation involved 857 real-world case studies. There, PH-LLM performed similarly to human experts on fitness-related tasks and improved over the base Gemini model in providing personalised sleep insights [s1]. That is a more modest claim, and a more meaningful one: parity with experts on open-ended cases rather than superiority on exams.

Finally, the model predicted self-reported sleep quality using a multimodal encoding of wearable sensor data, which the authors present as evidence it can contextualise sensor modalities rather than merely describe them [s1].

The paper releases its datasets, rubrics and benchmark performance to let other groups build on the work [s1] — a genuine contribution, and one that the second August paper suggests is the exception rather than the norm.

Why the second paper exists

The CHART Statement — Chatbot Assessment Reporting Tool — was published in JAMA Network Open on 1 August [s2]. Its stated rationale is blunt: the rise in chatbot health advice studies has been accompanied by heterogeneity in reporting standards, which affects how interpretable those studies are [s2].

CHART was developed through a comprehensive systematic review that identified variation in how such studies are conducted, reported and designed; a draft checklist was then revised through an international modified asynchronous Delphi consensus process involving 531 stakeholders, three synchronous panel consensus meetings of 48 stakeholders, and pilot testing [s2]. The result is 12 items and 39 subitems [s2].

The items are worth listing in outline, because they map the things that routinely go unreported: model identifiers, model details, prompt engineering, query strategy, performance evaluation, sample size, data analysis, disclosures, funding, ethics, protocol and data availability [s2].

The gap the checklist exposes

Consider what "model identifiers" and "prompt engineering" mean in practice. A study reporting that a chatbot answered health questions at some accuracy rate is describing a system consisting of a base model, a specific version of that model, a system prompt, a query format and a decoding configuration. Change any of those and the number changes. Commercial models are also updated without notice, so a result obtained in March may not reproduce in September on what is nominally the same product.

CHART's authors position the checklist as support for clinicians, researchers, editors, peer reviewers and readers in reporting, understanding and interpreting these findings [s2]. The word "readers" is doing real work there. Chatbot health advice studies get press coverage far out of proportion to their methodological maturity, and the headline number — an accuracy percentage, an exam score — is almost always the least transferable part of the result.

What neither paper claims

PH-LLM was evaluated on benchmarks and case studies. Neither the paper nor this account describes a trial in which people used it and their sleep or fitness measurably changed. The distinction between "the model produces advice experts rate as good" and "the advice makes people healthier" is the distinction that matters for anyone deciding whether a wearable's AI coach is worth attending to, and it remains unaddressed.

The CHART statement, for its part, is a reporting guideline. It standardises how studies are described; it does not make the underlying studies better designed, and it cannot make a chatbot's advice safe. Reporting checklists in other fields — CONSORT for trials, PRISMA for reviews — took years to change practice and never eliminated poor studies. They made poor studies easier to recognise.

What to watch

Whether journals adopt CHART as a submission requirement is the concrete near-term question, because that is the mechanism by which reporting guidelines actually bite. The second is whether the consumer-facing versions of models like PH-LLM are evaluated in real use rather than on benchmarks — and whether those evaluations, when they come, report the details CHART asks for.

Sources

Sources

  1. A personal health large language model for sleep and fitness coachingNature Medicine , August 14, 2025
  2. Reporting Guideline for Chatbot Health Advice Studies: The CHART StatementJAMA Network Open , August 1, 2025

More on

Related coverage