Most AI models that predict who will skip their medicines aren't ready for the clinic
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.
A systematic review published on 2 September in npj Digital Medicine delivers a blunt verdict on a fast-growing corner of medical AI: most published models that try to predict which patients will fail to take their medicines are not ready to be used in the clinic [s1].
The problem the models address is real and large. The World Health Organization has long estimated that adherence to long-term therapy for chronic illness averages only about 50% in developed countries, which makes reliably flagging patients at risk of stopping treatment a genuinely valuable target [s2]. The question the review asks is whether the AI tools built for that job actually work well enough to trust.
What the review examined
The authors reviewed 41 studies that developed models to predict medication non-adherence, and appraised each using PROBAST+AI — a structured tool for judging the quality and risk of bias of prediction models, covering the study's participants and data sources, its predictors, its outcome definitions and its analysis [s1].
Using a standardised appraisal tool matters here, because the failure modes of prediction models are often invisible in a paper's headline accuracy number. A model can report an impressive score and still be untrustworthy if it was trained and tested on the same narrow data, or if the thing it is predicting was loosely defined.
What it found
The results are unflattering. Most models showed "great concern" for development quality, at 71%, and a high risk of bias in how they were evaluated, at 80% [s1]. The recurring culprits were basic: poorly defined adherence outcomes, inadequate handling of missing data, and limited validation [s1].
Each of those culprits maps to a concrete way a tool can fail in the real world. When the outcome itself — whether someone "adhered" — is loosely defined, two studies can report similar accuracy while predicting different things, and neither transfers cleanly to a clinic that measures adherence its own way [s1]. Inadequate handling of missing data is not a technicality either: in adherence records, the fact that data are missing is often informative, and mishandling it can quietly inflate a model's apparent performance [s1]. Limited validation — testing a model only on the data it was built from — is the failure most likely to produce a number that looks good on paper and collapses on a new population [s1].
The most pointed finding concerns complexity. Reported discrimination — how well a model separates patients who will adhere from those who won't — did not consistently improve as the algorithms got more sophisticated [s1]. The authors' conclusion is that methodological rigour, not the choice of model type, is the key barrier to translating these tools into practice [s1]. In plain terms: a more elaborate neural network is not what stands between the field and a usable tool; careful study design is.
Why it matters
This is a useful corrective to a common assumption in medical AI — that adopting a more powerful model is the path to a more useful one. On the evidence of these 41 studies, the bottleneck lies upstream, in how the studies are built and checked [s1]. A model that predicts non-adherence poorly, but is deployed as if it predicts it well, can misdirect scarce follow-up resources or wrongly label patients, so the gap between reported and real-world performance is not academic.
The review does not claim that adherence prediction is hopeless, nor that no individual model is sound. It argues that the field's average quality is low and that the fixes are known. The authors offer framework-guided recommendations aimed at improving the robustness, interpretability and clinical actionability of future models [s1] — a to-do list, not an obituary.
For readers following the arrival of AI in everyday care, the takeaway is a habit of mind rather than a verdict on any one product: the right question about a predictive tool is not how advanced its algorithm is, but how carefully it was validated and on whom. On current evidence, most medication-adherence models have not cleared that bar. This article is informational and does not evaluate or recommend any specific tool.
Sources
- [s1] "AI models for medication adherence prediction: closing the gap to clinical readiness." npj Digital Medicine, 2 September 2026.
- [s2] World Health Organization, "Adherence to Long-Term Therapies: Evidence for Action," 2003.
Sources
- AI models for medication adherence prediction: closing the gap to clinical readiness — npj Digital Medicine , September 2, 2026
- Adherence to Long-Term Therapies: Evidence for Action — World Health Organization , January 1, 2003
When AI symptom-checkers disagree with patients, patients tend to walk away
In 6,772 users of a US health system's AI triage tool, people engaged about twice as often when the AI matched what they already planned — raising a hard question about what the tools are steering.
Clinical AI has 63 ways to measure fairness and one built for clinical use
A Lancet Digital Health scoping review found the field's fairness metrics fragmented and rarely clinically validated. A second review found that most studies don't measure fairness at all.
A computer that grades each colonoscopy raised how often endoscopists found adenomas
A Danish stepped-wedge trial gave endoscopists automated feedback on their technique after every procedure. Adenoma detection rose from 43.4% to 48.6% — a different tool from real-time polyp AI.
An ECG 'foundation model' matched rivals using a fraction of the labelled data
Trained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.