Research abstract · Draft
Quantifying Latent Malpractice Exposure in Ambient Clinical LLMs: An NCC MERP-Anchored Runtime Validation Framework
Ose Okodogbe, B.A. Molecular & Cell Biology (UC Berkeley) · PharmD/MBA Candidate, 2029
Background
Ambient clinical LLMs and AI scribes are being deployed faster than health systems can validate them. Hallucinated drugs, missed contraindications and dosing errors create latent malpractice exposure that existing vendor self-evaluations do not measure in clinically meaningful units.
Objective
To propose and evaluate CAIVI, an independent runtime validation framework that scores AI outputs with two indices — the Clinical Accuracy Index (CAI) and the Clinical Risk Index (CRI) — anchored to the NCC MERP medication-error harm taxonomy.
Methods
AI outputs are classified into six error classes (contraindication/allergy breach, wrong dose/route, major drug–drug interaction, hallucinated drug/indication, omission/guideline deviation, fabricated citation). Each class carries a severity weight, a probability of harm and a probability of clinician interception. Projected harm per output is Σ (error rate × P(harm) × (1 − catch rate × vigilance) × care-setting acuity). Error rates are reported with 95% Wilson score intervals. A planned retrospective pilot will compare 25–50 de-identified AI-generated notes against clinician-verified records, with pharmacist adjudication of each discrepancy.
Expected results
We hypothesise that CRI will separate tools with similar headline accuracy but different harm profiles, and that high-acuity settings (ICU/ED, pediatrics) will show disproportionate risk from dosing and contraindication errors.
Conclusions
Translating AI errors into NCC MERP harm bands and projected harm per 100,000 outputs gives risk managers, counsel and malpractice carriers a defensible, auditable metric for clinical AI governance.
Limitations
Severity weights, harm probabilities and catch rates are currently illustrative and require calibration against real adverse-event data. Results are not yet validated.
Keywords: clinical AI governance, large language models, medication safety, NCC MERP, pharmacovigilance, malpractice risk.