FDA opens comment on how to regulate AI medical devices that generate their own answers
A new discussion paper proposes a two-axis risk framework and a physician-training analogy for evaluating generative AI devices. It is not a rule, and the agency is asking what one should look like.
The Food and Drug Administration on August 18 issued a discussion paper asking the public how it should regulate medical devices built on generative AI — software that produces novel, unscripted output rather than a fixed classification or score — and opened a formal comment docket, FDA-2026-N-7874, that runs through October 19 [s1]. The document is explicitly not a proposed rule or draft guidance. It is a set of questions, with the agency's own tentative thinking attached, about how to evaluate a category of device that did not really exist for regulatory purposes until the last two years [s1][s2].
The distinction matters because the FDA has already cleared well over a thousand AI-enabled devices, but the overwhelming majority are narrow classifiers — software that flags a nodule on a chest CT or measures an ejection fraction from an echocardiogram. Those devices produce one of a limited set of predetermined outputs. A generative AI device can produce an open-ended answer, including one that synthesizes information the manufacturer did not explicitly program in. The FDA's Digital Health Center of Excellence, which is leading the effort, says such devices "hold transformative promise for patient care and the broader health ecosystem" while introducing "unique risks when compared to traditional software and other AI-enabled medical devices" [s1].
What the paper actually proposes
Three pieces of the discussion paper are doing the real work.
The first is a "two-axis" framework for assessing risk, intended to help regulators and manufacturers sort GenAI devices by how they're used and how much autonomy they're given, rather than treating all generative AI as a single risk category [s1]. The FDA has not published the finished axes as fixed policy — the paper lays out a "possible" framework and asks for reaction to it [s1].
The second is a premarket evaluation concept the agency describes as "inspired at a high level by how physicians are trained and evaluated": a competency assessment made up of non-clinical device benchmarking plus clinical confirmation, meant to establish that a GenAI device performs as intended before it reaches patients [s1]. That is a notable analogy for a regulator to reach for — it treats demonstrating competence, not just accuracy on a fixed test set, as the standard a generative device should clear.
The third is postmarket monitoring. Because generative models can drift, be updated, or behave differently outside the conditions they were evaluated in, the paper describes "several potential approaches to risk-proportionate postmarket monitoring" and raises the harder problem of foundation models and agentic AI systems — devices that don't just answer a question but take multi-step actions toward a goal [s1]. Both are areas where the FDA acknowledges it does not yet have a settled position, which is why the paper is structured as a series of targeted questions for each topic rather than a set of requirements [s1][s2].
Why the timing is not incidental
The discussion paper lands roughly eight months after the first device combining a large language model with agentic clinical functionality — an insulin-dosing support tool called UpDoc — cleared FDA review in December 2025 and began commercial deployment in June 2026 at health systems including Cleveland Clinic, Allegheny Health Network, and UCSF Health [s3]. That device illustrates the boundary the FDA is now trying to draw a framework around: its LLM layer holds a conversation with the patient and collects information, but the actual insulin-dosing logic is deterministic and provider-configured, keeping the generative component away from the decision itself [s3]. The discussion paper's premarket and postmarket questions are, in effect, asking what should be required of the next device that draws that line differently — one where the generative layer is closer to, or part of, the clinical decision.
Acting FDA Commissioner Kyle Diamantas framed the effort as a matter of global positioning as much as domestic policy: "Artificial intelligence is transforming medicine, and the United States must lead in shaping how this technology is developed and used safely and responsibly" [s1]. CDRH Director Michelle Tarver said the goal is a process that "safeguards patients and consumers, advances innovation, and serves as a potential model for regulators around the world" [s1]. Both framings point toward the same practical fact: no other major regulator has finished writing GenAI-specific device rules either, so whatever the FDA settles on will likely function as a reference point beyond U.S. borders.
What this does not do
The paper creates no new legal obligation for manufacturers and clears no new device. It does not resolve, on its own, how a generative device's premarket testing should differ from a standard classifier's, what an acceptable postmarket drift-monitoring plan looks like, or how liability and update cadence should work for a model that can be retrained after clearance. Those are precisely the questions the docket is collecting answers to before the FDA drafts anything binding [s1][s2].
What to watch
Whether the comment period, which closes October 19, produces a draft guidance in the following months, and whether the "two-axis" risk framework survives contact with manufacturer and clinician feedback in a recognizable form [s1]. Also worth watching: whether the agency's physician-training analogy for premarket evaluation — benchmarking plus clinical confirmation — becomes the template other regulators borrow, given the FDA's own stated ambition to set a model followed elsewhere [s1].
Sources
- [s1] FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices. U.S. Food and Drug Administration, August 18, 2026. https://www.fda.gov/news-events/press-announcements/fda-seeks-public-feedback-inform-regulatory-approach-generative-ai-enabled-medical-devices
- [s2] Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. U.S. Food and Drug Administration, Digital Health Center of Excellence, August 18, 2026. https://www.fda.gov/medical-devices/digital-health-center-excellence/considerations-regulation-generative-ai-enabled-medical-devices-discussion-paper-and-request
- [s3] First FDA-Cleared AI Agent and LLM Enabled Device Confirmed. Innolitics, June 25, 2026. https://innolitics.com/articles/updoc-fda-cleared-ai-agent/
Sources
- FDA Seeks Public Feedback to Inform Regulatory Approach for Generative AI-Enabled Medical Devices — U.S. Food and Drug Administration , August 18, 2026
- Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback — U.S. Food and Drug Administration, Digital Health Center of Excellence , August 18, 2026
- First FDA-Cleared AI Agent and LLM Enabled Device Confirmed — Innolitics , June 25, 2026
95.5% of FDA AI device summaries omit who the model was trained on
A census of 691 AI-enabled devices cleared through 2023 found demographic data missing almost everywhere, three summaries reporting patient outcomes, and 113 recalls driven mostly by software.
The FDA has cleared 1,357 AI medical devices. Three were tested on patient outcomes.
A researcher who expected the evidence base to be thin says even she was surprised by how thin. Most cleared devices never appear in a registered clinical trial at all.
FDA creates a new device category for AI that reads skin wounds without touching them
The agency granted its first authorization for a software-aided adjunctive diagnostic device in wound assessment, a category built for tools that analyze a wound optically.
FDA clears an AI tool that flags pulmonary hypertension risk from a standard ECG
The clearance is the third in a growing family of machine-learning notification algorithms built to spot serious cardiac conditions from a routine 12-lead ECG, a test most patients already get.