CAIVI

Research abstract · Draft

Quantifying Latent Malpractice Exposure in Ambient Clinical LLMs: An NCC MERP-Anchored Runtime Validation Framework

Ose Okodogbe, B.A. Molecular & Cell Biology (UC Berkeley) · PharmD/MBA Candidate, 2029

Background

Ambient clinical LLMs and AI scribes are being deployed faster than health systems can validate them. Hallucinated drugs, missed contraindications and dosing errors create latent malpractice exposure that existing vendor self-evaluations do not measure in clinically meaningful units.

Objective

To propose and evaluate CAIVI, an independent runtime validation framework that scores AI outputs with two indices — the Clinical Accuracy Index (CAI) and the Clinical Risk Index (CRI) — anchored to the NCC MERP medication-error harm taxonomy.

Methods

AI outputs are classified into six error classes (contraindication/allergy breach, wrong dose/route, major drug–drug interaction, hallucinated drug/indication, omission/guideline deviation, fabricated citation). Each class carries a severity weight, a probability of harm and a probability of clinician interception. Projected harm per output is Σ (error rate × P(harm) × (1 − catch rate × vigilance) × care-setting acuity). Error rates are reported with 95% Wilson score intervals. A planned retrospective pilot will compare 25–50 de-identified AI-generated notes against clinician-verified records, with pharmacist adjudication of each discrepancy.

Expected results

We hypothesise that CRI will separate tools with similar headline accuracy but different harm profiles, and that high-acuity settings (ICU/ED, pediatrics) will show disproportionate risk from dosing and contraindication errors.

Conclusions

Translating AI errors into NCC MERP harm bands and projected harm per 100,000 outputs gives risk managers, counsel and malpractice carriers a defensible, auditable metric for clinical AI governance.

Limitations

Severity weights, harm probabilities and catch rates are currently illustrative and require calibration against real adverse-event data. Results are not yet validated.

Keywords: clinical AI governance, large language models, medication safety, NCC MERP, pharmacovigilance, malpractice risk.