ANALYSIS

Experts rated chatbot sleep advice for athletes as valid. Jet lag advice mostly failed

Twenty experts scored three language models on ten common questions. One model cleared the content validity bar on both topics; another scored zero on jet lag.

Athletes and the staff around them ask chatbots practical questions — when to nap, how to handle a five-timezone flight before a competition. A study published on 3 October is one of the few attempts to have domain experts score what the models say back [s1].

The result splits cleanly by topic, and the split is informative about where language models are reliable.

The design

The study ran in two phases between January and June 2024 [s1]. First, ten frequently asked questions on athlete sleep and jet lag were identified with input from experts and from the models themselves [s1]. Second, 20 experts — mean age 43.9 ± 9.0 years, ten female and ten male — assessed the models' responses via surveys administered at two time points [s1].

Three models were tested: ChatGPT-3.5, ChatGPT-4 and Google Bard [s1]. Response rates were high, at 100% at the first time point and 95% at the second [s1].

Three statistics were computed: Fleiss' Kappa for agreement between raters, the Jaccard Similarity Index for each rater's consistency with themselves, and the content validity ratio (CVR) for whether responses were judged appropriate [s1].

The results

Content validity separated the models sharply. ChatGPT-4 had the highest CVR for sleep at 0.67, and was the only model with a valid CVR for jet lag, at 0.68 [s1]. Google Bard showed the lowest CVR for jet lag — 0% — with significant differences against ChatGPT-3.5 (p = 0.0073) and ChatGPT-4 (p < 0.0001) [s1].

Reasons for inappropriate responses varied significantly for jet lag (p < 0.0001), with Google Bard criticised for insufficient information and frequent errors [s1]. ChatGPT-4 outperformed the other models overall [s1].

The authors' conclusion is measured: the study highlights the potential of LLMs, particularly ChatGPT-4, to provide evidence-based advice on sleep, while underscoring the need for improved accuracy and validation for jet lag recommendations [s1].

Why jet lag is the harder problem

The topic split is not arbitrary, and it says something about what these models do well.

General sleep advice for athletes is largely stable, consensus-level content: sleep duration targets, sleep hygiene, nap timing, the effects of late-evening competition. It is well represented in training data and does not change much between individuals.

Jet lag advice is a computation. It depends on the direction of travel, the number of timezones crossed, the individual's chronotype, the departure and arrival times, the competition schedule, and the timing of light exposure and — where used — melatonin. Getting it right means applying circadian phase-response principles to a specific itinerary, and getting the direction wrong actively delays adaptation rather than merely failing to help.

A model that produces confident-sounding but directionally wrong phase-shift guidance is worse than one that declines. The CVR of 0% for one model on jet lag [s1] is what that failure mode looks like when it is scored.

The agreement problem

The reliability statistics deserve as much attention as the validity ones.

Inter-rater reliability was minimal, with Fleiss' Kappa between 0.21 and 0.39 [s1]. Intra-rater agreement was high, with 53% of experts achieving a Jaccard Similarity Index of 0.75 or above [s1].

That combination — experts consistent with themselves but not with each other — is a finding about the field, not about the models. It means there is no settled expert consensus on what a correct answer to these questions looks like. Each rater applied a stable internal standard; the standards differed.

This makes any single validity score harder to interpret. A CVR is computed against the panel's judgement, and if the panel does not agree with itself as a group, the threshold the models are being measured against is fuzzy. The study reports this openly rather than burying it.

What has already dated

The models tested were ChatGPT-3.5, ChatGPT-4 and Google Bard, evaluated between January and June 2024 [s1]. All three have since been superseded by later versions, and Bard has been rebranded.

This is the structural problem with evaluating commercial language models in peer-reviewed literature. The evaluation cycle takes 18 months; the product cycle takes three. A result about a specific model version is a historical record by the time it is published. What survives is the method and the topic-level finding — that a consensus-heavy topic scored well and a computation-heavy topic did not — rather than the model rankings.

The wider AI-in-sport picture

A narrative review published on 13 October surveys AI applications across endurance sport, covering metabolic health monitoring, recovery prediction and nutrition personalisation [s2]. It reports that AI systems integrate multimodal physiological, environmental and behavioural data, and that continuous glucose monitoring combined with AI algorithms allows carbohydrate management during prolonged events [s2].

Its limitations section is the relevant part here. The review identifies persistent challenges including measurement validity and reliability of sensor-derived signals, dataset quality problems such as noise, missingness and labelling error, model performance and generalisability, algorithmic transparency, and equitable access [s2]. It notes that limited generalisability due to homogeneous training datasets restricts applicability across diverse athletic populations [s2].

The authors undertook no formal risk-of-bias assessment or meta-analysis, citing heterogeneity [s2] — so the review maps a field rather than grading it.

What neither study establishes

Nothing here measures whether an athlete who followed a chatbot's advice slept better, adapted faster, or performed differently. Content validity is a judgement about a text, not an outcome.

Twenty raters and ten questions [s1] is a small evaluation, and the questions were partly generated by the models being tested [s1] — a design choice that risks favouring topics the models are comfortable with.

What to watch

The useful next design is prospective: athletes randomised to model-generated versus practitioner-generated travel and sleep plans, with actigraphy and performance endpoints. Until that exists, expert-rated text quality is the ceiling of what is known.

Sources

  • [s1] "Can We Trust Them?" An Expert Evaluation of Large Language Models to Provide Sleep and Jet Lag Recommendations for Athletes — Sports Medicine, published online 3 October 2025. https://doi.org/10.1007/s40279-025-02303-5
  • [s2] Artificial Intelligence in Endurance Sports: Metabolic, Recovery, and Nutritional Perspectives — Nutrients, 13 October 2025. Narrative review. https://doi.org/10.3390/nu17203209

Sources

  1. "Can We Trust Them?" An Expert Evaluation of Large Language Models to Provide Sleep and Jet Lag Recommendations for AthletesSports Medicine , October 3, 2025
  2. Artificial Intelligence in Endurance Sports: Metabolic, Recovery, and Nutritional PerspectivesNutrients , October 13, 2025
Related coverage