ANALYSIS

95.5% of FDA AI device summaries omit who the model was trained on

A census of 691 AI-enabled devices cleared through 2023 found demographic data missing almost everywhere, three summaries reporting patient outcomes, and 113 recalls driven mostly by software.

When the FDA clears a medical device, it publishes a decision summary describing what the device does and what evidence supported the decision. A study published this month in JAMA Health Forum read every one of those summaries for AI-enabled devices and counted what they contain [s1].

The census

The study is cross-sectional, linking FDA decision summaries and approvals databases with the Manufacturer and User Facility Device Experience (MAUDE) database of adverse event reports and the FDA Medical Device Recalls Database [s1]. It covers all AI and machine learning devices cleared by the FDA from September 1995 to July 2023, with the analysis run from October to November 2024 [s1].

That comes to 691 devices, of which 254 — 36.8% — were cleared in or after 2021 [s1]. More than a third of nearly three decades of AI device clearances happened in the final two and a half years of the window.

What the summaries leave out

The reporting gaps are the finding. Decision summaries frequently omitted the study design (323 devices, 46.7%), the training sample size (368, 53.3%), and demographic information about the data the device was built on (660, 95.5%) [s1].

That last figure is the one with the clearest downstream consequence. For 95.5% of these devices, a clinician reading the public summary cannot tell whose data the model learned from — which ages, which sexes, which racial or ethnic groups, which care settings. Model performance in medical imaging and risk prediction is known to vary across populations, and a summary that omits the training demographics offers no way to reason about where the device might underperform.

On evidence quality: 6 devices reported data from randomized clinical trials (1.6% as reported) and 53 (7.7%) from prospective studies [s1]. Fewer than 40% of premarket summaries contained data published in peer-reviewed journals (272, 39.4%) [s1].

On performance reporting: sensitivity was given for 166 devices (24.0%), specificity for 152 (22.0%), and patient outcomes for 3 (less than 1%) [s1].

Three devices out of 691 whose public summary reports a patient outcome. Everything else is reported, when it is reported at all, in terms of technical task performance.

On safety documentation: 195 summaries (28.2%) reported safety assessments, 344 (49.8%) reported adherence to international safety standards, and 42 (6.1%) described risks to health [s1].

What has gone wrong in the field

The postmarket side of the study draws on the FDA's own adverse event and recall records. Across the 691 devices, 489 adverse events were reported involving 36 devices (5.2%): 458 malfunctions, 30 injuries and 1 death [s1]. Forty devices (5.8%) were recalled, 113 times in total, primarily because of software issues [s1].

Those numbers are lower than the reporting gaps might lead one to expect — but MAUDE is a passive, voluntary reporting system, and underreporting in it is well documented. The recall count is more reliable, and its cause is worth noting: software, not hardware, is what fails in this device class.

The authors conclude that despite increasing clearance of AI/ML devices, standardised efficacy, safety and risk assessment by the FDA is lacking, and that dedicated regulatory pathways and postmarket surveillance of AI/ML safety events may be needed [s1].

The trial evidence, where it exists

A systematic review published later in the month gives a sense of how much randomised evidence exists in the single specialty where AI adoption is furthest along. Searching MEDLINE, Web of Science and the Cochrane Library from inception to November 2024, it found 11 randomised controlled trials evaluating machine learning models against traditional methods in cardiovascular care [s2].

Eleven trials. They were conducted between 2021 and 2024, and 81.2% were multicentre [s2]. Five (45.5%) reported improvements in clinical events, six (54.5%) showed enhanced diagnostic accuracy or early detection, and three (27.3%) demonstrated improved resource utilisation [s2].

The authors' own summary is two-sided: the review highlights AI's potential to improve early detection, diagnostic accuracy and resource efficiency, while stating that the limited number of RCTs indicates a need for more high-quality studies to validate effectiveness across clinical domains [s2].

The limits of both papers

The JAMA Health Forum study measures what is disclosed in public summaries, not what the FDA reviewed. Manufacturers submit more to the agency than appears in a decision summary, and a missing item in a summary is not proof the agency never saw it. What the study establishes is that the public record — the document a hospital or clinician can actually consult — is largely silent on training data, demographics and outcomes [s1].

It also covers clearances only through July 2023 and reflects databases as of late 2024 [s1]. Given that a third of the clearances came in the final stretch of that window, the composition of this device class is changing faster than any census of it can track.

The cardiovascular review's 11 trials cover one specialty, and its percentages are computed over a very small denominator [s2].

What to watch

Whether decision summaries begin routinely including training-set demographics. That is a documentation change rather than a regulatory-standard change, and it is the single item in this study with the widest gap — 95.5% — between what is reported and what a clinician would need to judge whether a model applies to the patient in front of them [s1].

Sources

Sources

  1. Benefit-Risk Reporting for FDA-Cleared Artificial Intelligence-Enabled Medical DevicesJAMA Health Forum , September 5, 2025
  2. Randomized Controlled Trials Evaluating Artificial Intelligence in Cardiovascular Care: A Systematic ReviewJACC. Advances , September 24, 2025
Related coverage