Fail Systems · Health 201

Healthcare is becoming automated. What happens when the system fails?

Hospitals are turning into cognitive composite systems — humans, AI models, agents, devices, and infrastructure working as one. When the power goes out, the network drops, or the model is wrong, lives depend on what happens next. Fail Systems is a framework for understanding, classifying, and designing against that failure.

Framework

A taxonomy of failure

Failure in automated healthcare is not one thing. The Fail Systems framework classifies it into distinct layers — each with its own likelihood, blast radius, and defense.

Power & infrastructure

Grid loss, UPS exhaustion, cooling failure, cloud-region outage. The layer everything else stands on.

Connectivity & data

Network partitions, EHR downtime, ransomware, corrupted or unavailable patient data at the moment of decision.

Devices & electronics

Monitors, pumps, ventilators, lab systems — firmware faults, update cascades, silent sensor drift.

Models & agents

Wrong outputs delivered confidently, agent loops acting on stale state, automation that fails without announcing it.

The human handoff

The seams. Skills that atrophied under automation, alarms nobody owns, the moment a human must take over and can't.

Cascades

Small faults that propagate across layers — one bad update, one dead datacenter, one wrong default, system-wide.

Design principle

Graceful degradation, planned in advance

A safe automated system is one that knows how to become a less automated system. Every clinical workflow should have a defined answer at each tier — before the failure, not during it.

0

Full automation

AI and agents operate the workflow; humans supervise by exception.

1

Assisted operation

Automation degraded or distrusted — humans decide, machines advise.

2

Manual operation

Electronics up, cognition down. Staff run the workflow on devices alone.

3

Analog fallback

Power or network gone. Paper, batteries, hand calculation, human judgment — and a plan for it that has actually been rehearsed.

In progress

What's coming

The Fail Systems framework

The full methodology: failure classes, severity scoring, and degradation-tier design for clinical systems. FMEA for the AI-hospital era.

Coming soon

Failure readiness assessment

A structured self-assessment for hospitals and health-tech builders: where would your system break first, and what happens then?

Coming soon

Case studies

Real incidents — outages, update cascades, EHR downtime, AI failures — analyzed through the framework, built on the AI Med Risk incident library.

Coming soon
If you are in crisis: this site is about system safety, not a source of help. In the US, call or text 988 (Suicide & Crisis Lifeline). Elsewhere, see findahelpline.com.