Tech & Medicine

Most AI models that predict who will skip their medicines aren't ready for the clinic

A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.

A systematic review published on 2 September in npj Digital Medicine delivers a blunt verdict on a fast-growing corner of medical AI: most published models that try to predict which patients will fail to take their medicines are not ready to be used in the clinic [s1].

The problem the models address is real and large. The World Health Organization has long estimated that adherence to long-term therapy for chronic illness averages only about 50% in developed countries, which makes reliably flagging patients at risk of stopping treatment a genuinely valuable target [s2]. The question the review asks is whether the AI tools built for that job actually work well enough to trust.

What the review examined

The authors reviewed 41 studies that developed models to predict medication non-adherence, and appraised each using PROBAST+AI — a structured tool for judging the quality and risk of bias of prediction models, covering the study's participants and data sources, its predictors, its outcome definitions and its analysis [s1].

Using a standardised appraisal tool matters here, because the failure modes of prediction models are often invisible in a paper's headline accuracy number. A model can report an impressive score and still be untrustworthy if it was trained and tested on the same narrow data, or if the thing it is predicting was loosely defined.

What it found

The results are unflattering. Most models showed "great concern" for development quality, at 71%, and a high risk of bias in how they were evaluated, at 80% [s1]. The recurring culprits were basic: poorly defined adherence outcomes, inadequate handling of missing data, and limited validation [s1].

Each of those culprits maps to a concrete way a tool can fail in the real world. When the outcome itself — whether someone "adhered" — is loosely defined, two studies can report similar accuracy while predicting different things, and neither transfers cleanly to a clinic that measures adherence its own way [s1]. Inadequate handling of missing data is not a technicality either: in adherence records, the fact that data are missing is often informative, and mishandling it can quietly inflate a model's apparent performance [s1]. Limited validation — testing a model only on the data it was built from — is the failure most likely to produce a number that looks good on paper and collapses on a new population [s1].

The most pointed finding concerns complexity. Reported discrimination — how well a model separates patients who will adhere from those who won't — did not consistently improve as the algorithms got more sophisticated [s1]. The authors' conclusion is that methodological rigour, not the choice of model type, is the key barrier to translating these tools into practice [s1]. In plain terms: a more elaborate neural network is not what stands between the field and a usable tool; careful study design is.

Why it matters

This is a useful corrective to a common assumption in medical AI — that adopting a more powerful model is the path to a more useful one. On the evidence of these 41 studies, the bottleneck lies upstream, in how the studies are built and checked [s1]. A model that predicts non-adherence poorly, but is deployed as if it predicts it well, can misdirect scarce follow-up resources or wrongly label patients, so the gap between reported and real-world performance is not academic.

The review does not claim that adherence prediction is hopeless, nor that no individual model is sound. It argues that the field's average quality is low and that the fixes are known. The authors offer framework-guided recommendations aimed at improving the robustness, interpretability and clinical actionability of future models [s1] — a to-do list, not an obituary.

For readers following the arrival of AI in everyday care, the takeaway is a habit of mind rather than a verdict on any one product: the right question about a predictive tool is not how advanced its algorithm is, but how carefully it was validated and on whom. On current evidence, most medication-adherence models have not cleared that bar. This article is informational and does not evaluate or recommend any specific tool.

Sources

  • [s1] "AI models for medication adherence prediction: closing the gap to clinical readiness." npj Digital Medicine, 2 September 2026.
  • [s2] World Health Organization, "Adherence to Long-Term Therapies: Evidence for Action," 2003.

Sources

  1. AI models for medication adherence prediction: closing the gap to clinical readinessnpj Digital Medicine , September 2, 2026
  2. Adherence to Long-Term Therapies: Evidence for ActionWorld Health Organization , January 1, 2003
Related coverage