Guide · Risk & compliance
Healthcare compliance AI tools: how risk teams should evaluate clinical AI
For hospital risk managers, compliance officers and AI governance committees. Not legal advice.
Why clinical AI is now a compliance question
Health systems are adopting AI scribes, patient-message drafting and decision support faster than they can test them. When one of these tools suggests a wrong dose, misses an allergy or invents a drug, the liability sits with the organization that deployed it, not only with the vendor.
That moves clinical AI out of IT procurement and into the work of risk, compliance and governance teams, who need evidence they can document and defend.
What regulators and accreditors are signalling
US rules now touch clinical AI from several directions. The FDA regulates some AI software as a medical device. ONC's HTI-1 rule sets transparency requirements for decision support delivered through certified health IT. HHS's Section 1557 rule requires covered entities to identify and reduce discrimination risk from patient care decision support tools.
None of these replaces local due diligence. Organizations are still expected to know how the tools they use perform in their own patients and settings. Check current guidance with your counsel, as these rules continue to change.
Questions to ask any clinical AI vendor
What error types were measured, and how were they defined? Was testing done by someone independent of the vendor? Are results reported with confidence intervals and sample sizes? How do results differ by care setting, such as ICU, emergency or pediatrics? What happens when the model is updated, and is it re-tested?
A single headline accuracy number answers none of these. Two tools with the same accuracy can carry very different harm if one tends to make dosing or contraindication errors.
Where an independent audit fits
CAIVI is an investigational framework for auditing clinical AI outputs after the fact rather than interrupting clinicians with alerts. It sorts errors into six classes, weights them by potential harm using NCC MERP harm categories, and reports a Clinical Accuracy Index (CAI) and Clinical Risk Index (CRI) with confidence intervals.
The intended uses are pre-purchase vendor comparison, periodic re-audits for governance committees, and evidence for risk and insurance review. CAIVI's weights and scores are currently illustrative and have not yet been validated against real outcome data.