Regulating a Moving Target – FDA Seeks Comments on Possible Framework for Regulation of GenAI
September 3, 2026On August 18, 2026, FDA’s Center for Devices and Radiological Health (CDRH) released Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback (Discussion Paper). The paper is not draft or final guidance; rather, it is intended to solicit feedback on the regulation of generative artificial intelligence (GenAI)-enabled medical devices. It poses 26 discussion questions in the categories of assessment of risk, a competency-based approach for premarket evaluation, postmarket monitoring, foundation model device master files, and considerations for agentic AI systems.
The Discussion Paper does not address whether these approaches fall within FDA’s existing legal authorities. While the postmarket control ideas may require new authorities, the design, development, and premarket control suggestions appear to fit within existing processes.
The Discussion Paper proposes a two-axis risk framework: one axis considers the activity performed by a GenAI function (non-directive, action directing, action taking with HCP supervision, and fully autonomous), while the other considers consequences of relying on an incorrect output. This mirrors the ISO 14971 approach, which evaluates probability of harm and severity of harm. As GenAI activity progresses toward full autonomy, oversight decreases and the probability of harm from incorrect outputs increases; as consequences move from limited to severe, so does the severity of potential harm.
CDRH is considering a competency-based approach for premarket evaluation—modeled on medical training, licensure, supervised practice, periodic reevaluation, and public reporting—consisting of non-clinical device benchmarking and clinical confirmation. The final user-facing device would be evaluated based on its intended use and proportionate to its risk.
Benchmarking is a non-clinical evaluation using well-defined, reusable tests and datasets to measure AI performance at a specific point in time, assessing whether the device demonstrates the clinical knowledge, analytic capabilities, safety behavior, and generalizability needed for its intended use.
Because non-clinical benchmarking may not capture user interaction, workflow, or patient population issues, CDRH is also considering clinical confirmation approaches: retrospective evaluation on real patient inputs, shadow deployment, standardized patient interactions, clinician adjudication, or prospective studies. These methods are reasonable and have precedent with other device types. As with other devices, the challenge will be sample size determination and statistical analysis—CDRH acknowledges this and seeks stakeholder input on endpoints and sample sizes.
Benchmarking and clinical confirmation essentially describe verification and validation. Whether the new terminology adds clarity or confusion—particularly when integrating GenAI into multi-function devices and documenting testing in the quality management system—remains to be seen. The types of testing, methods, and acceptable data sources may be more relevant considerations.
The Discussion Paper also discusses standards against which performance of GenAI-enabled devices should be measured. Predetermined acceptance criteria should be derived from design and development inputs, yet while there is acknowledgement that GenAI-enabled devices produce varied outputs and may undergo continuous adjustments, there is no discussion on how users should construct design and development inputs for GenAI devices that will form the basis of acceptance criteria. This foundational question arguably should precede discussions of risk and assessment approaches.
CDRH is considering whether to accept greater premarket uncertainty in exchange for enhanced postmarket monitoring. Under the Quality Management System Regulation (QMSR), verification after release is triggered by design changes, but GenAI may involve continuous changes or performance drift even without discrete modifications. The Discussion Paper suggests periodic benchmarking, sample-based clinician review, and performance degradation monitoring as potential postmarket approaches. Predetermined change control plans in premarket submissions could help, and QMSR updates may also be necessary. We note that this is not the first time CDRH has suggested that postmarket data may be able to fill gaps related to premarket uncertainty; in its guidance document, Factors to Consider When Making Benefit-Risk Determinations in Medical Device Premarket Approval and De Novo Classifications, CDRH indicated that postmarket data may allow device authorization where questions about device effectiveness remain. In our experience, CDRH has, in reality, been reluctant to adopt this trade-off. If it does so for GenAI-enabled devices perhaps it would adopt a similar approach to other device types.
The Discussion Paper is an important step in FDA’s engagement with GenAI challenges. While the proposed approaches build on familiar regulatory concepts, their application to GenAI will require careful consideration of these technologies’ inherent variability and evolving nature. Stakeholders should take advantage of the comment period to provide input on practical implementation challenges. Comments are due October 19, 2026.