ANALYSIS

Seven chatbots were tested on drinking advice. All were kind; all got something wrong

Clinicians rated empathy at 4.6 out of 5 and quality of information at 2.7. Every chatbot produced at least one piece of guidance judged inappropriate, overstated or inaccurate.

January is the month when the largest number of people in English-speaking countries try to change their drinking, and it is now also the month when a good number of them will ask a chatbot for help. Two studies published within five days of each other in late January examined what happens next — one testing what chatbots say, the other testing whether it is possible to know in advance who a digital intervention will work for.

What the chatbots said

A study in NEJM AI on 22 January evaluated seven publicly available chatbots, including both general-purpose tools and ones built specifically for behavioural health [s1]. The researchers built a fictional case and simulated seven days of interaction, using 25 prompts derived from real Reddit posts about alcohol misuse [s1].

Four clinicians independently rated each transcript across five domains — empathy, quality of information, usefulness, responsiveness and scope awareness — using an evaluation framework designed for chatbots [s1]. They also assessed whether the chatbot used stigmatising language, and whether it ever challenged the user rather than only validating what the user said [s1].

The pattern in the results is consistent and uncomfortable. Empathy scored highest across all chatbots, at a mean of 4.6 out of 5 [s1]. Quality of information scored lowest, at 2.7 [s1]. Overall mean performance varied widely between tools, from 2.1 (SD 1.1) to 4.5 (SD 0.8) [s1]. Chatbots built for behavioural health did not perform significantly better than general-purpose ones [s1].

Every chatbot produced at least one example of guidance the clinicians judged inappropriate, overstated or inaccurate [s1]. None used stigmatising or judgmental language, and all supported self-efficacy [s1].

That combination is the finding. These are tools that are reliably warm and unreliably correct, which is precisely the profile most likely to be trusted past the point where it should be. The authors' framing is measured: chatbots were perceived to vary widely, are generally strong on empathy, and have room for improvement in response quality [s1].

What the study cannot tell you

This was a simulation, not a trial. No real person with a drinking problem was helped or harmed in it, and there is no outcome measured beyond four clinicians' ratings of transcripts. The prompts were derived from real posts but the case was fictional. Seven chatbots at one moment in time, on tools that are updated continuously, is a snapshot. The study does not name a best chatbot, and the range between 2.1 and 4.5 means the category average tells an individual user very little.

The second question: who does a digital tool work for?

Digital alcohol interventions do work for some people and not others, and the field has been poor at telling them apart in advance. A paper in npj Digital Medicine on 27 January reports that prior attempts to predict individual intervention effectiveness have performed only slightly above chance — around 0.60 for both area under the curve and balanced accuracy [s2].

The authors combined data across several domains: psychological assessments, social network data, and neural responses to alcohol cues [s2]. They then made predictions before delivering a smartphone intervention based on psychological distancing to young adults, in two studies of 67 and 114 participants [s2].

Random forest models predicted individual differences in effectiveness in the first study (balanced accuracy 0.71, 95% CI 0.69 to 0.73; AUC 0.87, 95% CI 0.85 to 0.88) and replicated in the external test sample, though less strongly (balanced accuracy 0.68; AUC 0.68, 95% CI 0.54 to 0.82) [s2]. The authors compare this with a clinical-utility threshold of 0.67 balanced accuracy drawn from prior digital health work [s2].

The substantive finding underneath the modelling is simpler than the modelling. Interventions were most effective for participants who perceived their peers as moderate but frequent drinkers [s2]. The authors suggest peer drinking perceptions could serve as a low-burden indicator for identifying likely non-responders early [s2].

Read together

Both studies are small, both are early, and neither should change what anyone does this month. But they point at the same gap. The chatbot study shows that the current generation of conversational tools delivers the relational part of behavioural support well and the informational part poorly. The prediction study shows the field is only now developing ways to tell whether a given digital tool will help a given person at all.

A tool that is warm, always available, free, and wrong some of the time is not obviously worse than nothing — but the case that it is better has not been made in a trial with real participants and real drinking outcomes. That trial has not been run.

What to watch

Whether any behavioural health chatbot publishes an evaluation against clinical outcomes rather than transcript ratings, and whether the peer-perception signal identified in the npj study replicates in a larger, more diverse sample.

Sources

  1. [s1] Assessing Generative AI Chatbots for Alcohol Misuse Support: A Longitudinal Simulation Study. NEJM AI, 22 January 2026. https://doi.org/10.1056/AIcs2500676
  2. [s2] Predicting individual differences in digital alcohol intervention effectiveness through multimodal data. npj Digital Medicine, 27 January 2026. https://doi.org/10.1038/s41746-026-02356-4

Sources

  1. Assessing Generative AI Chatbots for Alcohol Misuse Support: A Longitudinal Simulation StudyNEJM AI , January 22, 2026
  2. Predicting individual differences in digital alcohol intervention effectiveness through multimodal datanpj Digital Medicine , January 27, 2026

More on

Related coverage