A decade into AI drug discovery, the clinical evidence is 'disappointingly limited'
A Nature Reviews Drug Discovery audit says the problem isn't the models. They were built to be validated rather than used, and benchmarked against the wrong thing.
Artificial intelligence in drug discovery has had roughly a decade of sustained investment, a steady stream of partnership announcements, and an equally steady stream of benchmark results. A Perspective published in Nature Reviews Drug Discovery on 7 August audits what came of it, and opens by defining the test it intends to apply: what matters in drug discovery is delivering safer and more efficacious medicines to patients faster [s1].
Against that test, the paper's verdict is that although a wide variety of AI methods have been developed, applied and benchmarked, evidence of their clinically relevant impact is so far disappointingly limited [s1].
Where the field actually stands
The framing is worth holding alongside a separate marker from December. A Nature news feature published on 9 December 2025 was still asking which drug would be the first designed by AI, and canvassing antibody candidates as the leading contenders [s2] — an article that would not be written if the question had an answer.
That is the shape of the gap. Not an absence of AI-derived molecules in development, but an absence of ones that have completed the journey the field's promises are denominated in.
Five reasons the Perspective gives
The paper's contribution is that it does not stop at the observation. It sets out candidate explanations, and they are structural rather than technical [s1].
Clinical translation was not the design target. The Perspective identifies an insufficient focus on clinical translation during model development [s1]. A model optimised to score well on a retrospective benchmark is a different artefact from one optimised to change what a project team does next, and the field has produced far more of the first kind.
Life science data is conditional. The paper points to difficulties with applying AI algorithms on conditional life science data [s1] — measurements that hold under the assay, cell line, dose range and protocol they were generated in, and do not straightforwardly transfer outside them. This is the property that makes biological data unlike the domains where large-scale machine learning has produced its most visible wins.
The problems were underspecified. The Perspective names insufficient problem definitions and the resulting underspecification of computational models for real-world use cases [s1]. A model built to answer a question that was never precisely posed cannot be evaluated against a real decision, because no one wrote down what decision it was for.
Technology push rather than science pull. The paper identifies this dynamic as a likely underlying factor [s1]. Capability arriving in search of an application produces different work from a bottleneck in search of a capability, and the field has been shaped substantially by the former.
Operationalisation takes years. The Perspective notes the substantial time required to turn technical capabilities into systems that are sufficiently scaled and accessible for users [s1] — the unglamorous distance between a method that works in a paper and infrastructure a discovery organisation can actually run on.
The recommendation that follows
The paper's headline recommendation is a change in how the field evaluates itself: benchmarking studies of AI tools in drug discovery need to move on from model validation and instead focus on the tools' ability to improve decision making [s1].
This is a more consequential proposal than it sounds. Model validation asks whether predictions match held-out data. Decision-making improvement asks whether a project that used the tool reached a better answer, or the same answer sooner, than one that did not — which requires a counterfactual, a comparison group, and an outcome measured in project terms rather than in AUC.
That kind of evaluation is expensive, slow, and commercially awkward, since the comparison group is a version of a programme run without the tool the vendor is selling. Almost none of it exists. It is also the only form of evidence that would settle the question the Perspective poses.
What this Perspective is
A Perspective in a review journal is an argued position by its authors, not a systematic review or a quantitative synthesis. It does not report a count of AI-derived candidates by development phase, a success-rate comparison against conventionally discovered compounds, or a statistical test of any of its five explanations. Its authority rests on the authors' reading of the field rather than on a reproducible method, and readers should weight it accordingly.
What makes it useful anyway is the specificity of the diagnosis. "AI has not delivered yet" is a familiar complaint. "AI has not delivered because models were benchmarked on validation rather than on decision quality, on data that does not transfer across conditions, against problems that were never fully specified" is a set of claims that can be acted on, and — more importantly — checked.
What to watch
Whether any group publishes a decision-quality benchmark of the kind the Perspective calls for, comparing programmes run with and without a given AI tool. Whether a compound with both an AI-identified target and an AI-designed molecule reaches approval, and how its timeline compares with conventional programmes. And whether the antibody candidates the December Nature feature identified as leading contenders [s2] progress, since antibodies are the modality where structural prediction has the strongest claim to having changed the work.
Sources
- [s1] Artificial intelligence in drug discovery - what it is, where we stand and the path forward. Nature Reviews Drug Discovery, published online 7 August 2026. https://doi.org/10.1038/s41573-026-01496-2
- [s2] What will be the first AI-designed drug? These disease-fighting antibodies are top contenders. Nature, 9 December 2025. https://doi.org/10.1038/d41586-025-03965-x
Sources
- Artificial intelligence in drug discovery - what it is, where we stand and the path forward — Nature Reviews Drug Discovery , August 7, 2026
- What will be the first AI-designed drug? These disease-fighting antibodies are top contenders — Nature , December 9, 2025
An ECG 'foundation model' matched rivals using a fraction of the labelled data
Trained on 1.7 million ECGs paired with clinicians' report text, ECG-CLIP reached the same accuracy as the best comparator with about 90% less training data — a bid at the field's labelling bottleneck.
Most AI models that predict who will skip their medicines aren't ready for the clinic
A review of 41 studies found that the great majority of AI medication-adherence prediction models carried high risk of bias, and that fancier algorithms did not reliably predict better.
Nvidia and Lilly are spending $1bn on AI drug discovery. No results are attached
The announced lab pairs robotic wet labs with computational models in a continuous loop. It is a serious bet on a real bottleneck — and, so far, entirely a forward-looking statement.
An AI-designed drug for lung fibrosis has its first phase 2a results
Rentosertib was safe over 12 weeks in 71 patients with idiopathic pulmonary fibrosis. A lung-function signal appeared at the highest dose, in a trial not designed to prove efficacy.