Model never worked as well locally as claimed
A model validated on the developer's data is deployed across many hospitals without independent local validation. Discrimination, calibration and alert burden at the new site differ from the claims, and nobody measures sensitivity because only the alerts that fire are visible.[2,3,4]
Warning signs
- Vendor performance figures with no external validation at a comparable site
- Alert rate tracked but not missed cases
- Clinicians describe the alert as noise
- Large variation in performance between sites running the same model
Seen inEpic Sepsis Model missed two-thirds of sepsis cases in external validation
Dataset shift after deployment
The patient mix, disease patterns, coding practice or upstream data feed changes, and the relationship the model learned no longer holds. Output keeps flowing with no error message. The model can over-alert (flooding staff) or under-alert (missing cases), and the change is often noticed first by frontline staff, not by monitoring.[5,6,7]
Warning signs
- Sudden change in alert volume or score distribution
- New disease, new population or new upstream system (lab analyzer, EHR build, coding change)
- Nursing complaints of overalerting
- Missing values, duplicate records or type mismatches in model inputs
Seen inSepsis model switched off after COVID-19 changed the patient mix
Biased proxy labels and hidden subgroup failure
The model is trained to predict something convenient (cost, prior treatment, a clinician's order) instead of the clinical outcome. Where access to care differs by group, the proxy encodes that difference, and the model under-serves the same patients the system already under-serves. Imaging models can also detect attributes like race that humans cannot see, so bias can enter without an obvious input variable.[8,9,10,11]
Warning signs
- Target variable is cost, utilization or a clinician action rather than health status
- No performance reported by race, sex, age or disability
- Subgroup performance never checked in local data
Seen inCare-management algorithm under-referred Black patients because it predicted cost, not illness
Agents acting on stale, garbled or poisoned state
An agent reads state (a medication list, a refill queue, its own memory) that is out of date, mis-transcribed or deliberately poisoned, then acts on it through tools with real permissions. Each step looks locally reasonable; the error is only visible in the result, such as duplicate or wrong refills. Memory and context poisoning let one bad input shape behaviour long after it arrived.[21,22,23,24]
Warning signs
- Agent can write to production systems without a confirmation step
- No readback of drug name and dose to a human
- Rising complaints about duplicate or unexpected actions
- Agent memory persists across sessions with no review
Seen inPharmacy AI phone agent garbled drug names and placed wrong and duplicate refills
Excessive agency, loops and false self-reports
Agents given broad permissions and autonomy ignore instructions, repeat steps, stop early or verify their own work incorrectly. When they fail they may also misreport what happened, for example claiming recovery is impossible. In a clinical setting that means an automated action nobody approved and an inaccurate account of the damage.[25,26,27,22]
Warning signs
- Agent holds write or delete permissions it does not need
- No human approval gate for high-impact actions
- No independent log of what the agent actually did
- Termination conditions not defined
Seen inCoding agent deleted a production database during a code freeze, then misreported recovery (non-clinical analogue)
Silent model updates and unreviewed deployment
Vendors retrain or swap models, or tools go live without regulatory review or documented validation, and the deploying organization is not told or does not re-test. Performance changes after an update look the same as normal operation. Transparency requirements that would expose this are themselves in flux.[28,24,29,15]
Warning signs
- Contract has no notice-of-change clause
- No re-validation step after vendor updates
- Model version not visible to users or in logs
- Unclear whether the tool is an FDA-regulated device
Seen inOntario auditor: all 20 approved AI scribes produced inaccurate notes in testing