FDA Is Rethinking How GenAI Medical Devices Should Be Evaluated

A medical device powered by generative AI creates a validation problem that traditional software does not handle neatly. If the same clinical question can be asked many ways, and the system can produce several differently worded but acceptable answers, what exactly should a manufacturer test?

FDA is now asking that question too. In August 2026, the agency’s Center for Devices and Radiological Health released a discussion paper on GenAI-enabled medical devices. It explores risk assessment, premarket evidence, postmarket monitoring, foundation models, and agentic systems.

Important: This is a discussion paper, not a regulation, draft guidance, or statement of final FDA expectations. The competency concept remains exploratory.

The validation challenge

Traditional medical-device software is typically developed around defined requirements and expected behavior. Manufacturers select representative inputs, define expected results, and verify that the software meets those requirements. The test set cannot cover every real-world scenario, but the functional boundaries are usually clear enough for representative testing to be meaningful.

GenAI expands those boundaries. A clinician may phrase the same request in several ways, and the device may produce different answers each time. Multiple outputs may be clinically acceptable; another may be largely correct but contain one misleading statement. Exact-output verification therefore becomes less useful.

FDA acknowledges that it may be unreasonable to evaluate every conceivable input for devices with open-ended inputs and outputs. The question begins to shift from “Did the system produce the predetermined answer?” to “Can the finished device reliably perform its intended job within defined clinical and safety limits?”

FDA’s risk framing: activity and consequence

FDA presents a possible two-axis heuristic. One axis considers what the device does and how independently it does it, ranging from providing non-directive information to directing or taking action. The other considers the severity of harm if a user relies on an incorrect output.

A tool that drafts a clinical summary for physician review does not create the same risk as a tool that independently recommends a treatment change, even if both use the same foundation model. The practical regulatory questions are therefore not simply how sophisticated the model is, but:

·       What authority does the finished device have in the clinical workflow?

·       How much independent human review exists?

·       How strongly does the output direct a user toward an action?

·       What could happen if the output is wrong?

Competency-based evaluation

The paper’s most interesting idea is a competency-based approach to premarket evaluation. FDA draws a high-level analogy to clinician evaluation: clinicians are not tested against every scenario they may encounter, but through structured assessment, supervised practice, and continued evaluation. The analogy concerns evaluation methodology; it does not equate an AI system with a physician.

FDA outlines two broad components:

Non-clinical device benchmarking. The device could be evaluated for clinical knowledge, analytic capability, safety behavior, communication, and generalizability across a sufficiently challenging range of inputs.

Clinical confirmation. Evidence would assess whether benchmarked capabilities translate into appropriate performance for the intended population, users, workflow, and conditions of use.

Competency does not eliminate validation. It changes what validation may need to prove. Manufacturers would still need defensible acceptance criteria, representative test conditions, clinically appropriate comparators, subgroup analysis, and evidence proportionate to risk.

Defining competency may be harder than measuring accuracy

For an interpretive device, what counts as acceptable performance? A correct conclusion with a flawed explanation may still be unsafe. An accurate answer may include distracting or overly directive information. Strong average performance may conceal poor performance for a clinically relevant subgroup.

Before competency can be measured, it has to be defined. Teams may need criteria covering clinical correctness, reasoning quality, communication, safe refusal, uncertainty handling, generalizability, and performance across intended populations and settings. That front-loaded definition work may be more demanding, not less, than conventional pass-or-fail testing.

The regulatory focus is the finished device

FDA’s proposed unit of evaluation is the final user-facing device in the configuration intended for deployment, not the foundation model standing alone. That distinction matters. The same model could support clinical documentation, diagnostic assistance, patient communication, or treatment recommendations, yet those products could have very different intended uses, safeguards, and evidence needs.

A strong model benchmark does not establish that the medical device built around it is safe and effective. The application, interface, system prompts, retrieval methods, data sources, workflow integration, intended user, guardrails, and human-oversight controls all shape real-world behavior. The regulated product is the complete device system.

Intended use must be enforced in the product

Conversational systems may technically be capable of much more than the manufacturer intends. Consider a device intended to explain laboratory findings. A user asks, “Which medication should I start?” The model may be able to answer, but treatment recommendation may sit outside the device’s intended use.

The manufacturer must decide whether the device should refuse, redirect, provide only general information, or ask for clinician review. It must then show that the boundary holds across varied phrasing and contexts. For GenAI devices, intended use may need to be more than language in labeling. It may need to be actively and reliably enforced by product behavior.

Third-party foundation models and change control

Many manufacturers will build on foundation models controlled by third parties. They may control the medical application, prompts, retrieval layer, interface, additional training, and risk controls, but not the underlying model’s training data, architecture, updates, content policies, refusal behavior, or retirement schedule.

An upstream update can change clinically relevant behavior outside the manufacturer’s normal development process. Yet the device manufacturer remains responsible for the regulated product. Supplier oversight, version traceability, technical change detection, contractual notification, re-benchmarking, and regulatory change assessment may all become central lifecycle controls.

FDA is also asking how predetermined change control plan concepts might apply when future changes cannot be fully specified, and whether a voluntary Foundation Model Device Master File could give review teams confidential model information. Neither concept removes the sponsor’s responsibility to demonstrate the safety and effectiveness of its own finished device.

Postmarket monitoring becomes part of the evidence strategy

Premarket testing establishes performance at a point in time. After deployment, the model or surrounding architecture may change, users may develop new prompt patterns, clinical practice may evolve, and uncommon failure modes may emerge only after extensive use.

FDA is considering risk-proportionate monitoring approaches such as periodic re-benchmarking, sample-based clinician review, and monitoring for performance degradation. Practical signals could include:

·       increasing rates of clinically inappropriate or unsupported responses;

·       failure clusters associated with particular prompts, workflows, sites, or populations;

·       repeated attempts to push the device beyond its intended use;

·       changes in output patterns after an upstream model or architecture update; and

·       threshold breaches that trigger investigation, CAPA, revalidation, or regulatory assessment.

These are not only AI-engineering metrics. They may become quality-system and regulatory evidence used to show continued competency across the total product lifecycle.

Key questions regulatory teams should ask now

1.     Can we define both sides of the intended-use boundary? - State what the device must do and what it must not do.

2.     Which competencies are essential to safety and effectiveness? - Translate the clinical role into measurable capabilities and safety behaviors.

3.     What makes a response acceptable? - Address variable wording, incomplete explanations, unsupported claims, and overly directive output.

4.     Does the evidence represent real use? - Cover relevant users, prompt diversity, languages, clinical settings, and patient subgroups.

5.     How will out-of-scope requests be handled? - Test refusal, redirection, escalation, and human-oversight pathways.

6.     How dependent are we on a third-party model? - Map upstream controls, notification rights, versioning, and fallback plans.

7.     What changes could affect the validated state? - Define triggers for re-benchmarking, revalidation, CAPA, and regulatory review.

8.     Can we investigate a failure? - Retain enough traceability to reconstruct clinically significant behavior while protecting privacy.

9.     Which indicators continue after launch? - Set monitoring cadence, thresholds, ownership, and response protocols before deployment.

Conclusion

FDA has not decided what a GenAI medical-device framework will ultimately look like. But the discussion paper makes the underlying problem clear: a regulatory model built around predictable software behavior becomes harder to apply when a device can produce an enormous range of context-dependent outputs and evolve after deployment.

Competency-based evaluation is one possible answer. If the concept develops further, manufacturers may need to think less about proving that a system can generate one predetermined answer and more about demonstrating that the finished device consistently has the capabilities, boundaries, and safeguards required for its intended clinical role.

For regulatory teams, that shift would reach well beyond the submission. It could reshape intended-use design, verification and validation, risk management, supplier controls, change assessment, and postmarket monitoring across the product lifecycle.

References

1.     U.S. Food and Drug Administration. “Considerations for the Regulation of Generative AI-Enabled Medical Devices: Discussion Paper and Request for Feedback.” August 2026. FDA discussion-paper webpage

2.     U.S. Food and Drug Administration. “Considerations for the Regulation of Generative AI-Enabled Medical Devices.” Discussion paper (PDF), August 2026. Download the FDA PDF

3.     U.S. Food and Drug Administration. “Measuring and Evaluating Artificial Intelligence-Enabled Medical Device Performance in the Real-World.” Request for Public Comment, 2025. FDA real-world performance page

4.     Regulatory Affairs Professionals Society. “FDA seeks feedback on framework for regulating GenAI devices.” August 2026. RAPS article

Editorial note: This article is educational analysis of an exploratory FDA discussion paper and is not legal or regulatory advice.

Next
Next

SnowPro Core Certification Guide