2,000 audited outputs · Inpatient ward · clinician vigilance ×1.0
CAI · Clinical Accuracy Index
97.9
CRI · Clinical Risk Index
32.1 Elevated
| Category | MERP | Count | Share of harm |
|---|---|---|---|
| Omission / guideline deviation | D–E | 22 | 28% |
| Wrong dose / route | F–H | 9 | 23% |
| Hallucinated drug / indication | E–F | 14 | 20% |
| Major drug–drug interaction | E–G | 7 | 15% |
| Contraindication / allergy breach | G–I | 3 | 8% |
| Fabricated evidence / citation | B–C | 31 | 6% |
Recommendation
Prioritise remediation of omission / guideline deviation, which accounts for 28% of projected patient harm. Deploy runtime interception before expanding use in higher-acuity settings.
Illustrative model. Harm = Σ error rate × P(harm) × (1 − clinician catch rate) × care-setting acuity. Severity bands follow NCC MERP. Not a substitute for clinical validation.
The independent risk rating for medical AI.
CAIVI audits clinical AI tools before and after purchase, scoring accuracy and patient-harm risk so risk officers, P&T committees, and malpractice insurers can approve, price, and defend every deployment.
Independent
No vendor conflict of interest
Auditable
Logged evidence for every output
Underwritable
Structured risk data for carriers
For CFOs & risk officers
Estimate avoided harm events, malpractice exposure and insurance savings per year.
Institution profile
Adjust assumptionsAI outputs / year
720,000
Harm events avoided
100.8
Claims avoided
15.12
Liability avoided
$18.1M
Premium savings
$210k
CAIVI cost / year
$240k
Net annual value
$18.1M
75.5× return on spend
Illustrative estimate. Assumes 15% of harm events become claims and up to a 5% premium discount. Pricing is an example.
For legal & compliance
AI vendors can't impartially grade their own tools. CAIVI sits between the model and the clinician as a neutral auditor, giving counsel a defensible record.
ONC HTI-1
Transparency & risk management for decision-support tools
Every AI output is scored and logged with source attributes and a CAI / CRI record.
FDA SaMD / CDS guidance
Clinicians can independently review the basis of a recommendation
Flags unsupported claims and phantom citations; shows the patient-record fact behind each block.
EU AI Act (high-risk)
Human oversight, logging, post-market monitoring
Pass / warn / block decisions keep a clinician in the loop with a full audit trail.
Joint Commission & CMS
Patient-safety event reporting and prevention
Errors mapped to NCC MERP harm bands for direct use in safety reviews.
Malpractice carriers
Evidence of reasonable risk controls on AI use
Independent third-party validation report supports underwriting and defense.
Not legal advice. CAIVI is positioned as an advisory audit layer; final clinical decisions stay with licensed clinicians.
Assessment workspace
Import an existing assessment
Paste an audit or incident report, or upload a PDF. CAIVI reads it and sets every factor below.
Care setting
Errors found, by root cause
Missed allergy, renal/hepatic or pregnancy flag · MERP G–I
Unit, weight-based or organ-function miscalculation · MERP F–H
Ignored CYP, QTc or serotonergic interaction · MERP E–G
Fabricated fact presented as clinical truth · MERP E–F
Missing therapy, monitoring or follow-up · MERP D–E
Phantom study or misattributed source · MERP B–C
97.9
32.1
Error rate
4.30%
95% CI 3.50–5.28%
Harm / 100k
646
reach patient & harm
P(≥1 harm) / 1k
99.8%
per 1,000 encounters
What drives the risk
Fix first: omission / guideline deviation — 28% of projected patient harm.
Illustrative model. Harm = Σ error rate × P(harm) × (1 − clinician catch rate) × care-setting acuity. Severity bands follow NCC MERP.
Ground-truth check
Paste the verified note or correct treatment on the left, and what the AI scribe or assistant produced on the right. CAIVI finds every difference and scores it.
Runtime intercept simulator
FHIR bundle in
{
"resourceType": "Bundle",
"entry": [
{
"Patient": {
"age": 72,
"weight_kg": 68
}
},
{
"Condition": {
"code": "N17.9",
"display": "Acute kidney injury"
}
},
{
"Observation": {
"code": "Creatinine",
"value": "2.9 mg/dL",
"trend": "↑ from 1.1 in 48h"
}
},
{
"Observation": {
"code": "eGFR",
"value": 21
}
},
{
"MedicationStatement": [
"Furosemide 40 mg IV daily"
]
}
]
}Model draft
Start gentamicin 7 mg/kg IV q24h for suspected gram-negative sepsis.
Validation pipeline
5 resources · 1 active condition
drug=gentamicin · dose=7 mg/kg · interval=q24h
eGFR 21 + rising creatinine → aminoglycoside nephrotoxicity
Loop diuretic + aminoglycoside → ototoxicity risk
CRI 88 · MERP G–H
Audit packet out
Validating…
Simulated cases and timings for demonstration only — not clinical advice.
Benchmark
| Tool | CAI | CRI | Error rate | Harm / 100k |
|---|---|---|---|---|
| Vendor A · Ambient scribe | 98.6 | 22.8 Elevated | 2.50% | 432 |
| Vendor B · CDS chatbot | 98.0 | 30.3 Elevated | 3.88% | 603 |
| Vendor C · Med rec assistant | 99.0 | 19.6 Low | 1.64% | 363 |
| Vendor B + CAIVIsafest | 99.6 | 7.1 Low | 0.68% | 122 |
Sample vendor data (5,000 audited outputs each), not real benchmarks.
01
Every model output passes through CAIVI before reaching the clinician.
02
Claims are checked against guidelines, formularies and evidence at runtime.
03
Errors roll up into CAI and CRI — auditable numbers for procurement and insurers.
Ose Okodogbe · Founder
Mission: bring pharmacovigilance discipline — NCC MERP harm grading, interaction and dosing checks — and health-system risk economics to generative clinical AI.