ANALYSIS

An AI-designed drug for lung fibrosis has its first phase 2a results

Rentosertib was safe over 12 weeks in 71 patients with idiopathic pulmonary fibrosis. A lung-function signal appeared at the highest dose, in a trial not designed to prove efficacy.

Patients with at least one treatment-emergent adverse event over 12 weeksRentosertib 30 mg once daily: 72.2%; Rentosertib 30 mg twice daily: 83.3%; Rentosertib 60 mg once daily: 83.3%; Placebo: 70.6%0%45%90%Rentosertib 30 mg once daily72.2%Rentosertib 30 mg twice daily83.3%Rentosertib 60 mg once daily83.3%Placebo70.6%
Patients with at least one treatment-emergent adverse event over 12 weeks
GroupValue (%)
Rentosertib 30 mg once daily72.2
Rentosertib 30 mg twice daily83.3
Rentosertib 60 mg once daily83.3
Placebo70.6
Patients with at least one treatment-emergent adverse event over 12 weeks The trial's primary endpoint; 18 patients per active arm and 17 on placebo. Source: Nature Medicine

The claim that artificial intelligence will design new medicines has been made for roughly a decade. The number of AI-discovered or AI-designed drugs that have actually reached human trials remains, in the authors' own words, few [s1]. On 3 June, Nature Medicine published results from one of them.

The trial

Rentosertib, formerly ISM001-055, is a small-molecule inhibitor of TNIK. Both the molecule and the target were arrived at using generative AI — the target through AI-driven target discovery, the compound through generative chemistry [s1]. The paper describes it as first-in-class on both counts, and TNIK as a first-in-class target in idiopathic pulmonary fibrosis [s1].

IPF is an age-related progressive lung disease for which, as the authors note, no current therapy reverses the degenerative course [s1].

The study was a phase 2a multicentre, double-blind, randomised, placebo-controlled trial [s1]. Patients received 12 weeks of treatment in one of four arms: 30 mg rentosertib once daily (n=18), 30 mg twice daily (n=18), 60 mg once daily (n=18), or placebo (n=17) [s1]. It is registered as NCT05938920 [s1].

The primary endpoint was safety, and it was met

The primary endpoint was the percentage of patients with at least one treatment-emergent adverse event — a safety endpoint, not an efficacy one. Rates were similar across arms: 72.2% (13 of 18) on 30 mg once daily, 83.3% (15 of 18) on 30 mg twice daily, 83.3% (15 of 18) on 60 mg once daily, and 70.6% (12 of 17) on placebo [s1].

Treatment-related serious adverse event rates were low and comparable across groups [s1]. The most common events leading to treatment discontinuation were related to liver toxicity or diarrhoea [s1]. That is worth noting rather than skipping past: hepatic signals in a first-in-class compound are the kind of finding that larger and longer trials exist to characterise.

The lung-function signal

Secondary endpoints included pharmacokinetics, forced vital capacity, diffusion capacity for carbon monoxide, forced expiratory volume in one second, Leicester Cough Questionnaire score, six-minute walk distance, and the number and duration of hospitalisations for acute exacerbations [s1].

The result being widely noticed is forced vital capacity at the highest dose. Mean change in the 60 mg once-daily group was +98.4 ml (95% CI 10.9 to 185.9), against -20.3 ml (95% CI -116.1 to 75.6) in the placebo group [s1].

Several things about that number deserve stating plainly. It comes from 18 patients over 12 weeks. Its 95% confidence interval runs from 10.9 to 185.9 ml — a seventeen-fold range, with the lower bound barely above zero. It is one of a long list of secondary endpoints in a trial powered for safety. The authors' own conclusion is calibrated accordingly: that targeting TNIK with rentosertib is safe and well tolerated, and warrants further investigation in larger-scale trials of longer duration [s1].

What "AI-discovered" does and does not mean here

The interesting question is what part of the pipeline the AI actually replaced. On the account in the paper, generative AI was used to identify TNIK as a target in IPF and to generate the molecule [s1]. Everything downstream — preclinical work, dose selection, trial conduct, the 12-week randomisation, the safety monitoring — is conventional drug development, and the result is a conventional phase 2a result: a safety signal that permits a phase 2b, and an efficacy hint that does not establish anything.

That is not a criticism of the achievement. Getting a novel target and a novel molecule into a randomised human trial is the hard part, and doing it faster is the whole proposition. But the trial does not test whether AI-derived molecules work better than human-derived ones, because no such comparison exists in the design. It tests one compound against placebo in 71 patients.

A second June paper on what AI systems do in medicine

A different kind of AI result appeared three days later. In Nature Cancer, researchers described an autonomous clinical AI agent built on GPT-4 and equipped with precision-oncology tools: vision transformers for detecting microsatellite instability and KRAS and BRAF mutations from histopathology slides, MedSAM for radiological image segmentation, and web-based search tools including OncoKB, PubMed and Google [s2].

Evaluated on 20 realistic multimodal patient cases, the agent selected appropriate tools with 87.5% accuracy, reached correct clinical conclusions in 91.0% of cases, and accurately cited relevant oncology guidelines 75.5% of the time [s2]. Against GPT-4 alone, decision-making accuracy rose from 30.3% to 87.2% [s2].

Twenty cases is a small evaluation set, and "realistic" cases are not patients. The authors present the work as a foundation for deploying such systems, not as evidence that one should be deployed [s2].

Read together, the two papers mark the same boundary from opposite sides. Both show AI systems performing a task that used to require expert human labour. Neither shows a patient outcome. The distance between those two things is where the next several years of this field will be spent.

What to watch

Whether rentosertib enters a larger, longer trial with forced vital capacity as a pre-specified primary endpoint, and how the liver-related discontinuations behave at scale [s1]. And whether autonomous clinical AI agents are ever evaluated prospectively against real patients rather than case vignettes [s2].

Sources

Sources

  1. A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trialNature Medicine , June 3, 2025
  2. Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncologyNature Cancer , June 6, 2025
Related coverage