top of page

Governing AI in Healthcare: Why Smart Governance Unlocks Productivity

  • 31 minutes ago
  • 7 min read


The uncomfortable pattern in the 2026 data: the health systems getting the most productivity out of AI are not the ones deploying the fastest. They are the ones who can tell you what their tools are actually doing.


The gap nobody wants to name

Adoption is no longer the story. Roughly three in four U.S. health systems have implemented or are planning at least one AI solution, and half of those running AI are running three or more applications at once. Among organizations that could quantify returns, more than half reported at least a 2x ROI — with AI-based clinical documentation improvement and denial prediction leading at roughly 70% hitting 2x or better, and ambient listening close behind at 61% (Eliciting Insights, Health System Adoption of AI Solutions 2026).

Then look at the operating model underneath.

In a Q2–Q3 2026 poll of 507 healthcare executives and informatics leaders, only 18.1% cleared a basic threshold of a governance committee, defined impact metrics, and post-deployment monitoring. Formal AI literacy training reached 18.3% of respondents. Meanwhile 67.7% were live, piloting, or actively evaluating ambient AI (Black Book Research, AI in the EHR Adoption Readiness Poll).


A separate 2026 survey found 94% of healthcare respondents claiming governance controls are in place — but 13% conceded those controls are applied only after issues surface. That is not governance. That is incident response wearing a governance badge.

The gap between "we have AI" and "we can account for our AI" is now the single largest predictor of whether the productivity thesis holds.


The productivity evidence is real — and narrower than the sales deck

This is where governance stops being a compliance argument and becomes an operations argument.

The strongest evidence for ambient documentation concerns clinician experience, not throughput:

Read those four findings together and the governance case writes itself.


Two products marketed identically produced a 9.5% gain and a rounding error in the same institution, in the same trial. The variance is not noise to be averaged away — it is the entire decision. An organization without local measurement has no way to know which of those two outcomes it purchased.

And the throughput finding matters for how you build the business case. If you promised your CFO more visits per session and the mechanism is actually reduced cognitive load and retention, you will report a failure against the wrong denominator and kill a program that was working.


What ungoverned deployment actually costs

The productivity downside of weak governance isn't a fine. It's rework, alert fatigue, and clinician trust — all of which are productivity line items.

Documentation errors are not rare at the pilot stage. In a pragmatic prospective pilot at UC Davis, 31 physicians generated 7,545 ambient notes; of 356 physician-evaluated notes, accidental omissions appeared in 18%, hallucinations in 11.5%, and accidental inclusions in 9.3%. Vendor-reported hallucination rates cluster at 1–3%; independent evaluations have landed considerably higher, with physical exam sections repeatedly identified as the highest-risk area — systems have documented entire examinations that never occurred.


Two things follow. First, 1–3% is not a comforting number when multiplied across millions of encounters and applied to medications, allergies, and diagnoses. Second, no ambient vendor accepts clinical liability for generated notes — the signing clinician does. That makes edit-rate monitoring and error taxonomy a clinical safety function, not an IT metric.

The canonical failure is still the most instructive one. The Epic Sepsis Model was deployed at hundreds of U.S. hospitals before Michigan Medicine externally validated it across 38,455 hospitalizations. Discrimination came in at AUC 0.63, against the 0.76–0.83 the developer had cited. Sensitivity was 33%; positive predictive value 12%. Clinicians would have needed to work through roughly 109 alerts to find one sepsis case they hadn't already caught (Wong et al., JAMA Internal Medicine, 2021; see also the accompanying editorial, The Epic Sepsis Model Falls Short — The Importance of External Validation). A 2024 external validation across two county emergency departments found sensitivity of 14.7% and PPV of 7.6% at the vendor-recommended threshold.


Epic subsequently overhauled the model and began recommending that hospitals train it on their own patient data before deployment. That recommendation is the lesson: local validation is not bureaucratic drag, it is the only thing standing between a purchased model and a decade of alert fatigue.

Alert fatigue is a productivity tax that compounds quietly. It never shows up as a governance failure on a dashboard. It shows up as clinicians ignoring the tool.


The regulatory floor is moving — in both directions at once

The strategic reason to build governance capability now is that the external scaffolding is becoming less reliable, not more.

Federal transparency requirements are being rolled back. HTI-1 established the first meaningful federal AI transparency floor in healthcare, requiring certified health IT developers to publish 31 source attributes for predictive decision support interventions — effectively model cards — so buyers could judge whether a tool was fair, appropriate, valid, effective, and safe. In December 2025, ASTP/ONC proposed HTI-5, which would remove the model card requirements entirely, on the reasoning that the agency found no publicly available evidence the requirements improved patient care and that clinicians rarely accessed source attribute information in workflow. Comments closed February 27, 2026; no final rule has issued. The AHA has urged the agency to retain the DSI criterion.

Whatever your view of the merits, the strategic implication is unambiguous: if the federal transparency floor drops, the burden of knowing what a model does shifts entirely onto the buyer. Governance stops being a compliance overlay and becomes a procurement capability.

State requirements are moving the opposite way. Utilization review is the clearest front. California's SB 1120 barred AI from being the sole basis for denying, delaying, or modifying care on medical necessity grounds, reserving final determinations for licensed clinicians. Texas SB 815, Maryland, Nebraska, Arizona, and Connecticut followed; 2026 added Washington, Iowa, Indiana, and Alabama, with Utah and Georgia effective in 2027. California's AB 3030 separately requires disclosure when generative AI produces clinical patient communications.


Device and international rules keep tightening. FDA's January 2025 draft guidance on AI-enabled device software functions set total product lifecycle expectations — data lineage, bias analysis, human-AI workflow description, validation tied to claims, and post-market performance monitoring. Predetermined Change Control Plans are now the mechanism for shipping model updates without a new submission, which rewards organizations that already monitor performance. In the EU, AI-enabled devices must satisfy MDR/IVDR and the AI Act, whose Article 50 transparency obligations began applying on 2 August 2026 to any provider whose outputs are intended for use in the EU.

The through-line: fragmentation is the permanent condition. A health system operating across three states, buying from six vendors, cannot compliance-chase its way through this. It needs one internal operating model that produces evidence any of these regimes will accept.


What "smart governance" actually means

The Joint Commission and CHAI published the first national guidance on responsible AI use in healthcare in September 2025, organized around seven areas including AI policies, local validation, ongoing monitoring, transparency, and education — explicitly designed to be adaptable to organizations at any maturity level. CHAI followed with governance playbooks in May 2026 aimed at making the framework operational, including for community health centers and safety-net settings.

Strip it to the four moves that pay for themselves:

1. One accountable clinical owner per tool. One in four organizations is attempting AI adoption with nobody accountable for the strategy. A pilot without an owner has nowhere to go once initial enthusiasm fades — it just quietly persists, unmeasured. Clinical leadership owns outcomes; IT owns integration and security; executives own funding.

2. Local validation before enterprise deployment. Vendor-reported performance is a hypothesis about your population, not a finding. The Epic sepsis case is what happens when that hypothesis goes untested at scale. Ninety-two percent of surveyed leaders now say deep clinical domain expertise is critical in evaluating a vendor — a direct lesson from the last cycle.

3. Post-deployment monitoring with defined thresholds. Drift is not an exception; it's the default. Monitoring is also increasingly the thing regulators and PCCP pathways assume you already have. Track edit rates, error taxonomy by severity, alert-to-action ratios, and override patterns — not just satisfaction scores.

4. Metrics tied to the actual mechanism. If the benefit is cognitive load and retention, measure burnout, after-hours time, and turnover intent. Do not promise throughput the evidence does not support.


The reframe

The dominant industry framing treats governance as friction — the thing that slows deployment down. The 2026 data supports the opposite reading.

Ungoverned AI produces pilots that never scale, tools clinicians learn to ignore, business cases measured against the wrong outcome, and error rates nobody discovers until they are in the permanent medical record. Every one of those is a productivity loss, and none of them appear on a compliance report.

Governed AI produces the one thing that actually compounds: the ability to tell, quickly and credibly, which of your deployments is working — and to kill or scale accordingly.

Adoption hit 75%. Governance is at 18%. The next eighteen months belong to organizations that close their own gap before a regulator, a plaintiff, or a burned-out medical staff closes it for them.

 
 
 

Comments


bottom of page