Sentir

FDA Asked What Postmarket Monitoring Should Look Like for Generative AI. We Read All 95 AI PCCPs Before Answering.

By Sentir Health · 09/15/2026

FDA's discussion paper on generative AI-enabled medical devices is open for public comment under docket FDA-2026-N-7874. It is the agency's first extended attempt to say how a device built on a large language model should be evaluated before authorization and watched after it. Postmarket monitoring is one of its central threads. Several of the discussion questions ask, in effect, whether the agency can accept more uncertainty at the front door if the monitoring behind it is good enough, and what "good enough" would have to mean.

We filed a comment on September 12. It is now posted on the docket, and a PDF copy is here. This post is the short version, and the story of what we read before we wrote it.

We read the record first

The paper asks how cadence and triggering events for postmarket reassessment should be set, whether re-benchmarking, clinician review, and degradation monitoring are the right tools, and whether machine-based supervisory agents can help. Those are questions about practice, so we went to the only public record of practice that exists: the authorized Predetermined Change Control Plans in the PCCP Tracker.

As of August 30, 2026, the tracker held 95 AI/ML-enabled devices authorized with a PCCP, with decision dates from February 2020 through July 2026. We pulled the full text of every public record for those devices, the 510(k) summaries, De Novo decision summaries, and clearance letters, and read every passage that touched change control or monitoring against a fixed rubric.

Two caveats govern everything that follows, and we put both in the comment itself. First, these are almost entirely non-generative devices. We did not find an authorized generative AI device with a publicly described PCCP. So the record is precedent from the adjacent category, the monitoring practice the current framework has actually produced for the device class nearest to generative AI. It says nothing about generative AI devices themselves. Second, public summaries are abridged. The detailed plan, with its protocols and thresholds, lives in the full submission, which is not public. Eight of the 95 records confirm a PCCP exists and describe none of its contents. Everything below is about what the public record discloses, not about what the underlying plans contain.

What the record shows

The postmarket evaluation these summaries describe is almost entirely change-triggered. A sponsor decides to modify the model, tests the modification against prespecified acceptance criteria, locks it, and releases it. That loop is well built. Quantitative acceptance criteria for modifications are routine, and many summaries disclose them down to the confidence-interval margin.

Monitoring of the fielded device, between changes, is a different story. Roughly one in five records invokes it in any form, and most of those do so in a sentence: a reference to complaint handling, to "post market surveillance" whose content is not described, or to real-world feedback as a reason for future retraining. Only a handful describe an actual method.

Across all 95 records we found exactly one specified monitoring cadence, Velmeni's quarterly evaluation, and exactly one quantified trigger for action, iSchemaView's retraining trigger at a stated sensitivity or specificity drift. We found no record describing scheduled re-benchmarking of an unchanged deployed model, and none describing postmarket monitoring at the subgroup level.

To be clear about what that does and does not mean. Silence in a public summary is not evidence that a company is not watching its model. In our experience nearly every team is, with scripts, dashboards, and someone's monthly notebook. What the record shows is narrower and, we think, more important: almost none of that watching is committed to in a form anyone outside the company could verify. The discipline that governs how these devices change has not yet been extended to how they are watched. Data collected to no standard is telemetry, not evidence.

What we told FDA

Our comment answers six of the discussion questions, the ones where our engineering work and the record above give us standing. The positions reduce to a few points.

A shift from premarket to postmarket evidence is sound only under four conditions. The monitoring program has to be continuous from the first day of deployment, because the intervening period cannot be reconstructed once a signal appears. Its analyses, thresholds, sampling frames, cadence, and triggers have to be prespecified before deployment, for the same reason FDA already requires prespecification of premarket benchmarks. Its records have to be independently verifiable: append-only, timestamped, and held under controls where selective retention or retrospective revision would be detectable. And it needs a prespecified consequence pathway, because monitoring without defined consequences is observation, not control.

The three proposed evaluation approaches are layers, not a menu. Re-benchmarking gives you comparability back to the authorization baseline but is blind between runs. Clinician review gives you depth but cannot scale to interaction volume. Degradation monitoring gives you coverage and timeliness but watches proxies, and a proxy can hold flat while quality falls. Any one alone inherits its own blind spot in full.

Cadence should be event-driven with a calendar floor. The events are prespecifiable: an input-distribution shift beyond a stated bound, any change to the device or a dependency, a deployment-volume milestone. The floor exists because a degradation monitor that never alarms is consistent with both a stable device and a broken monitor, and only a scheduled re-benchmark tells the two apart.

If machines help do the monitoring, hold the monitor to the same standard as the device. A supervisory agent should have characterized error rates against human adjudication, at the thresholds it runs at. It should be structurally independent from the device it watches, including in what model it is built on, because a monitor sharing the device's blind spots overstates protection exactly where it matters. Its outputs should be logged append-only. And any change to the supervisor, model version, prompt, threshold, or sampling logic, is a change to the monitoring program and should be versioned and re-validated as one.

Generative AI systems change through components, not retraining events. Every modification category in the current record is built around the model artifact and its training data. A generative AI device can change materially through a prompt edit, a retrieval-corpus update, a guardrail change, or a third-party model version that ships on someone else's schedule. We suggested that PCCPs for these devices categorize modifications by the component changed, with a prespecified mapping from each change class to the benchmark elements it can affect. That is the existing PCCP discipline, extended to the parts a generative AI system is actually made of.

Third-party model changes need detection, not just notification. A platform's changelog describes intent, not effect on your device. Version pinning is the strongest single mitigation and still leaves a window. A fixed, versioned canary set run on a cadence against the production model is the only mechanism that catches a silent upstream change before your clinical outputs do.

The bar everywhere else in software

Nobody buys enterprise software on "trust us, we do security." Buyers ask for a SOC 2 report: independent, tamper-evident, continuously renewed proof. We require that from the vendor that stores our files. There is no equivalent yet for the model reading a patient's scan, and the record above is what its absence looks like in public documents. The discussion paper is the first place FDA has asked, in writing, what that equivalent should be. That is why we filed.

A note to fellow builders

If your regulator opens a comment period in your lane, file. Ours took a weekend. It is the most leverage a small company ever gets over the rules it will live under, and the agency reads the docket.

If you hold a PCCP, are filing one, or have one on the roadmap, we would like to talk through the monitoring regimen behind it. You can see where your device sits, and who else in your clinical panel has committed to what, in the Sentir PCCP Tracker, browsable by category, clinical panel, PCCP type, or company.


Sentir Health is the independent system of record for the performance of FDA-cleared AI. We keep the baseline and the evidence your PCCP requires you to produce on the day you exercise it. Learn more or book a call.

Methodology note: The review covers the 95 AI/ML-enabled devices flagged as PCCP-authorized in the Sentir PCCP Tracker as of 2026-08-30, after excluding two records whose own documents state that no PCCP was included. Sources are public FDA 510(k) summaries, De Novo decision summaries, and clearance letters. Monitoring-related passages were classified manually against a fixed rubric; the strongest claims (one cadence, one quantified trigger, no subgroup-level monitoring) were re-verified by keyword sweeps over the full text of all 95 documents. Category and monitoring labels are editorial classifications based on public documents, not official FDA designations. The full comment, including the questions answered, is posted on regulations.gov and here as a PDF.