Key Takeaways
Overview
On August 18, 2026, the FDA’s Center for Devices and Radiological Health (CDRH) published Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback. The paper provides one of the FDA’s most detailed articulations to date of how the agency may approach regulation of GenAI-enabled medical devices. Among other things, it has two notable features: (1) a function-specific, two-dimensional framework for evaluating risk, and (2) a competency-based assessment approach that incorporates benchmarking and clinical confirmation in the premarket setting, along with ongoing monitoring in the postmarket setting.
The discussion paper is limited to the FDA’s regulation and oversight of medical devices enabled by GenAI and does not address the broader use of AI in clinical care. It is not draft guidance, does not propose or establish FDA policy, and acknowledges that the paper does not address whether the approaches under discussion fall within its existing statutory authority or would require additional congressional authorization.
The FDA is seeking comments through October 19, 2026. Stakeholders interested in clinical AI may wish to review the discussion paper and consider responding to the agency’s questions. Although the FDA emphasizes that the paper does not represent agency policy, it provides an important window to influence the direction of future policymaking.
Framework for Evaluating Risk
The FDA presents the following diagram as a possible two-axis heuristic for assessing the risk of GenAI-enabled software functions. Risk increases based on the degree and independence of the device’s activity and the severity of harm that could result from reliance on an incorrect output. By design, the chart does not establish device status, risk classification, or binding evidentiary requirements. However, it likely provides insight into the FDA’s current thinking regarding device classification and the potential boundaries of enforcement discretion.

The horizontal axis tracks the activity performed by the function and the degree of independence with which it operates. The vertical axis reflects the severity of harm that could occur if a user relied on an incorrect output. In other words, an action-taking function does not automatically fall into the highest-risk category. Its risk profile still depends on the potential consequences of an incorrect action and the degree of professional oversight involved.
The FDA also identifies factors that may shift where a function falls on the matrix. For example, a measurement or signal-processing function may provide non-directive information yet still present significant risk if users cannot independently assess whether the output is correct. Likewise, the same output may have different risk implications depending on whether it is delivered to a patient or a healthcare professional, or to a generalist rather than a specialist.
Competency-Based Assessment
The paper’s central discussion framework is a competency-based approach inspired, at a high level, by how clinicians are evaluated and credentialed. Importantly, the thing to be assessed is the final user-facing device as configured and intended for real-world deployment, rather than a foundation model standing alone.
Under the approach being considered, premarket competency would have two parts: non-clinical device benchmarking and clinical confirmation. Postmarket monitoring could then assess whether the device continues to perform against that baseline as use conditions and technical components change.
Benchmarking
The FDA uses “benchmarking” to mean well-defined, published, and reusable tests and datasets used to measure and compare performance at a particular point in time. The agency also contemplates that sponsor-developed tests tailored to a specific device and intended use could be used.
Overall, the paper organizes benchmarking into three baseline competency domains (safety, clinical proficiency, and generalizability), with an additional conditional domain for agentic devices, as shown in the diagram below:

Safety asks whether the device avoids dangerous behavior. Clinical proficiency asks whether it can perform its intended task. Generalizability asks whether those capabilities persist across the range of conditions the device is likely to encounter in practice. For agentic devices, the FDA presents additional competencies rather than a fourth baseline domain applicable to all GenAI-enabled devices.
Overall, benchmark results for a foundation model alone are unlikely to establish the competency of a downstream medical device. A device’s performance may be shaped by factors such as prompting, retrieval, guardrails, workflow design, and the user interface. As a result, the FDA appears focused on evaluating the deployed device as a whole rather than relying solely on evidence generated at the foundation model level.
Clinical Confirmation
Benchmarking, however comprehensive, may not fully capture how a GenAI-enabled device performs when used by actual users in real-world workflows and patient populations. CDRH is therefore considering clinical confirmation as an additional step to evaluate whether the configured device performs as intended in actual or clinically representative conditions of use. Clinical confirmation would not necessarily require a prospective clinical study in every case. Rather, the approach and level of evidence could be tailored to the device’s intended use, data modality, and risk profile.

Note, however, that the FDA is also asking stakeholders to help determine how that proportionality should work. Important design questions remain unresolved. The FDA asks how endpoints and sample sizes should be selected when clinical confirmation is not centered on traditional effectiveness measures. For open-ended outputs, potential comparators could include a clinical reference standard, a qualified clinician panel reflecting the standard of care, or clinician performance representative of real-world practice. The FDA also notes that, depending on the intended use, the relevant unit of analysis may be the human-AI team rather than the AI system alone. Finally, the agency asks whether performance should sometimes be evaluated against the care likely to occur in the absence of the device, such as unaided clinical judgment, delayed specialist review, or no intervention.
Postmarket Monitoring
CDRH is considering whether, in some circumstances, greater premarket uncertainty regarding benefit and risk could be justified by increased reliance on postmarket monitoring. The paper identifies three monitoring approaches:
The FDA also asks whether machine-based supervisory agents could facilitate monitoring and, if so, how the supervisory agent itself should be evaluated.
Other Considerations
Conclusion
The FDA has sought input on generative AI before, but this is the clearest indication yet of how the agency may actually regulate these technologies. Many of the questions the FDA poses reflect a practical effort to adapt existing device regulatory concepts to the realities of GenAI-enabled products. Stakeholders that want to influence where the agency ultimately lands should take the opportunity to engage now, while these ideas are still being debated and refined. Comments are due October 19, 2026.
Wilson Sonsini routinely advises clients on health AI governance, regulatory strategy, and regulator engagement, including evaluating existing processes against emerging regulatory expectations and preparing comments to federal agencies. For more information, please contact Wilson Sonsini attorneys Jodi Daniel and Ty Kayam, or any member of Wilson Sonsini’s Healthcare and FDA Regulatory practice.