Independent clinical AI audit & underwriting

CAIVI

The independent risk rating for medical AI.

CAIVI audits clinical AI tools before and after purchase, scoring accuracy and patient-harm risk so risk officers, P&T committees, and malpractice insurers can approve, price, and defend every deployment.

01

Independent

No vendor conflict of interest

02

Auditable

Logged evidence for every output

03

Underwritable

Structured risk data for carriers

For CFOs & risk officers

Model your clinical AI exposure

Estimate avoided harm events, malpractice exposure and insurance savings per year.

Institution profile

Adjust assumptions

AI outputs / year

720,000

Harm events avoided

100.8

Claims avoided

15.12

Liability avoided

$18.1M

Premium savings

$210k

CAIVI cost / year

$240k

Net annual value

$18.1M

75.5× return on spend

Illustrative estimate. Assumes 15% of harm events become claims and up to a 5% premium discount. Pricing is an example.

For legal & compliance

An independent safety record for every AI output

AI vendors can't impartially grade their own tools. CAIVI sits between the model and the clinician as a neutral auditor, giving counsel a defensible record.

ONC HTI-1

Transparency & risk management for decision-support tools

Every AI output is scored and logged with source attributes and a CAI / CRI record.

FDA SaMD / CDS guidance

Clinicians can independently review the basis of a recommendation

Flags unsupported claims and phantom citations; shows the patient-record fact behind each block.

EU AI Act (high-risk)

Human oversight, logging, post-market monitoring

Pass / warn / block decisions keep a clinician in the loop with a full audit trail.

Joint Commission & CMS

Patient-safety event reporting and prevention

Errors mapped to NCC MERP harm bands for direct use in safety reviews.

Malpractice carriers

Evidence of reasonable risk controls on AI use

Independent third-party validation report supports underwriting and defense.

Not legal advice. CAIVI is positioned as an advisory audit layer; final clinical decisions stay with licensed clinicians.

Assessment workspace

Score a clinical tool

Import an existing assessment

Paste an audit or incident report, or upload a PDF. CAIVI reads it and sets every factor below.

Care setting

Errors found, by root cause

Missed allergy, renal/hepatic or pregnancy flag · MERP G–I

Unit, weight-based or organ-function miscalculation · MERP F–H

Ignored CYP, QTc or serotonergic interaction · MERP E–G

Fabricated fact presented as clinical truth · MERP E–F

Missing therapy, monitoring or follow-up · MERP D–E

Phantom study or misattributed source · MERP B–C

CAI · Clinical Accuracy

97.9

CRI · Clinical RiskElevated

32.1

Error rate

4.30%

95% CI 3.50–5.28%

Harm / 100k

646

reach patient & harm

P(≥1 harm) / 1k

99.8%

per 1,000 encounters

What drives the risk

Omission / guideline deviation28%
Wrong dose / route23%
Hallucinated drug / indication20%
Major drug–drug interaction15%
Contraindication / allergy breach8%
Fabricated evidence / citation6%

Fix first: omission / guideline deviation — 28% of projected patient harm.

Illustrative model. Harm = Σ error rate × P(harm) × (1 − clinician catch rate) × care-setting acuity. Severity bands follow NCC MERP.

Ground-truth check

Compare AI output to the real record

Paste the verified note or correct treatment on the left, and what the AI scribe or assistant produced on the right. CAIVI finds every difference and scores it.

Runtime intercept simulator

Pick a patient. Watch CAIVI stop the error.

FHIR bundle in

{
  "resourceType": "Bundle",
  "entry": [
    {
      "Patient": {
        "age": 72,
        "weight_kg": 68
      }
    },
    {
      "Condition": {
        "code": "N17.9",
        "display": "Acute kidney injury"
      }
    },
    {
      "Observation": {
        "code": "Creatinine",
        "value": "2.9 mg/dL",
        "trend": "↑ from 1.1 in 48h"
      }
    },
    {
      "Observation": {
        "code": "eGFR",
        "value": 21
      }
    },
    {
      "MedicationStatement": [
        "Furosemide 40 mg IV daily"
      ]
    }
  ]
}

Model draft

Start gentamicin 7 mg/kg IV q24h for suspected gram-negative sepsis.

Validation pipeline

  1. ✓Parse FHIR bundle

    5 resources · 1 active condition

  2. ✓Extract claims

    drug=gentamicin · dose=7 mg/kg · interval=q24h

  3. ✕Renal check

    eGFR 21 + rising creatinine → aminoglycoside nephrotoxicity

  4. ✕Interaction check

    Loop diuretic + aminoglycoside → ototoxicity risk

  5. ✓Score

    CRI 88 · MERP G–H

Audit packet out

Validating…

Simulated cases and timings for demonstration only — not clinical advice.

Benchmark

Compare vendors side by side

ToolCAICRIError rateHarm / 100k
Vendor A · Ambient scribe98.622.8 Elevated2.50%432
Vendor B · CDS chatbot98.030.3 Elevated3.88%603
Vendor C · Med rec assistant99.019.6 Low1.64%363
Vendor B + CAIVIsafest99.67.1 Low0.68%122

Sample vendor data (5,000 audited outputs each), not real benchmarks.

01

Intercept

Every model output passes through CAIVI before reaching the clinician.

02

Validate

Claims are checked against guidelines, formularies and evidence at runtime.

03

Quantify

Errors roll up into CAI and CRI — auditable numbers for procurement and insurers.

Clinical origins

Built from the pharmacy side of patient safety

Read the research abstract

Ose Okodogbe · Founder

  • MCB · B.A. Molecular & Cell Biology, UC Berkeley
  • CRC · Clinical Research Coordinator Trainee, UC Davis Health (GCP, IRB)
  • PharmD/MBA · Candidate, class of 2029

Mission: bring pharmacovigilance discipline — NCC MERP harm grading, interaction and dosing checks — and health-system risk economics to generative clinical AI.