AWS us-east-1 DynamoDB DNS failure disrupts cloud-hosted clinical systems
A race condition in DynamoDB's DNS automation broke a core AWS region for about 15 hours. Some cloud-hosted EHR users slowed or went to paper; others saw nothing.[22,23,24]
how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
Not a layer: the way one failure crosses power, connectivity, devices, models and handoff, and spreads from one organisation to many.
Cascades is not a sixth layer. It is a propagation pattern that runs across the other five: power, connectivity, devices, models and handoff. A cascade starts as a failure in one layer, at one organisation, and turns into a failure in other layers or other organisations. Examples: a remote-access portal is breached and national claims processing stops. A security-software update crashes Windows machines and hospital services go offline. A pathology supplier is hit and three hospital trusts have to call for O-type blood donors.
Safety science has argued for decades that serious accidents in complex systems do not come from one broken part. Perrow called accidents in systems that are both interactively complex and tightly coupled 'normal accidents': they are to be expected, not freak events. Reason's Swiss-cheese model describes harm reaching a patient only when latent conditions and active failures in several defensive layers line up. Leveson's STAMP and CAST treat accidents as a loss of control over the system, not a chain of failed parts. Cook and Rasmussen described how hospitals chasing efficiency 'go solid': the buffers that used to soak up a problem disappear, so that an event in one distant part of the hospital suddenly matters everywhere else. None of these models is settled: professionals disagree on what the parts of the Swiss-cheese model mean, and its critics call it too linear. We use them as lenses, not as laws.
On this site a cascade is described by its path (the order in which it crossed layers, such as devices → connectivity → handoff) and its reach (one department, one hospital, a region, a country). Tracing the path matters because the controls that stop a cascade usually sit at the boundaries between layers and between organisations, and those boundaries are the parts nobody owns.
FailSystems viewAutomation does not add many new ways for a single part to fail. What it adds is coupling. When a hospital runs on a shared EHR, a shared clearinghouse, a shared endpoint agent and a shared pathology network, the same fault reaches every place at once, and the paper workaround has usually withered from lack of use. In our judgement the incidents that do the most harm to patients in automated care will be cascades. Most of them will start outside the hospital, in a supplier the hospital does not control and may not know it depends on. Plan for common-mode failure, not for one component failing on its own.
One software or content update is pushed to every installation at once and carries a latent defect that testing missed. Because every machine runs the same code, redundancy inside the hospital does not help: the primary and the backup crash together. Updates designed to ship fast, like security content, are the most exposed.[5,6,7]
Warning signs
Seen inCrowdStrike Falcon content update crashes Windows hosts, including hospital systems
Efficiency work strips out slack: spare beds, spare stock, spare staff, manual steps. Activities then depend directly on events elsewhere in the system, so a delay in one place becomes a stoppage in another within hours. In a tightly coupled process there is no time to improvise before the next step needs the output of the failed one.[8,9,10]
Warning signs
Seen inSynnovis pathology ransomware, South-East London, Texas winter storm: record load shed, hospitals lose water and heat
Hospitals depend on utilities that depend on each other. Loss of electricity can stop water treatment and gas supply, which in turn stops hospital heating, sterilisation and toilets. Patients at home on powered equipment lose it at the same time and arrive at the emergency department. The hospital's generator covers its own electricity but not the water pressure or the patients' home equipment.[11,12,13,14]
Warning signs
Seen inTexas winter storm: record load shed, hospitals lose water and heat
A hospital that goes to downtime diverts ambulances and patients to its neighbours. The neighbours have their own systems intact but not the extra capacity, so waits, walk-outs and delays in time-critical care rise there too. One organisation's cyber incident becomes a capacity incident for the region.[15,16,17]
Warning signs
Seen inRansomware spillover to adjacent San Diego emergency departments, WannaCry ransomware across the NHS in England
Cutting network links to contain an attack is often the right call, but it causes a cascade of its own. Partners that relied on the link lose the service, and organisations that were never infected shut systems down as a precaution because they lack clear central advice. The outage from containment can be larger than the outage from the attack.[1,16]
Warning signs
Seen inChange Healthcare ransomware and national claims/pharmacy clearinghouse outage, WannaCry ransomware across the NHS in England
Most cascades need several existing weaknesses at once: an unpatched system, a portal without multi-factor authentication, a missing bounds check, an untested backup. Each one is tolerable alone and may sit unnoticed for months. The cascade happens when a trigger finds a path through all of them. Use the Swiss-cheese picture with care. A survey of quality and safety professionals found they read its parts (holes, slices, arrow) in very different ways, and critics argue it is too static and linear. Its defenders still consider it useful because it is systemic.[18,16,1,5,19,20,21]
Warning signs
Seen inWannaCry ransomware across the NHS in England, Change Healthcare ransomware and national claims/pharmacy clearinghouse outage, CrowdStrike Falcon content update crashes Windows hosts, including hospital systems
A race condition in DynamoDB's DNS automation broke a core AWS region for about 15 hours. Some cloud-hosted EHR users slowed or went to paper; others saw nothing.[22,23,24]
A grid collapse cut power to continental Spain and Portugal for about ten hours. Hospitals largely held on generators; care outside them did not.[25,26,27,28,29]
A faulty Rapid Response Content update to CrowdStrike's Falcon sensor crashed about 8.5 million Windows devices worldwide. Outside-in measurement found disrupted services at 759 of 2,232 US hospitals studied.[5,30,31,6,7,32,33]
PathDevices → Connectivity & data → Human handoff
Ransomware hit Synnovis, the pathology provider for several south-east London NHS trusts and GP practices. Blood testing and matching collapsed, more than 11,000 appointments and procedures were postponed, O-type blood ran short nationally, and one death was later partly attributed to a delayed result.[34,35,36,37,10,38,39,3]
PathConnectivity & data → Human handoff
A ransomware attack took Ascension's electronic records offline for about five weeks. Clinicians told KFF Health News of medication errors and delayed lab results, and one said he had no training for the attack; Ascension said its care teams were trained for such disruptions.[40,41]
Attackers used stolen credentials on a Change Healthcare Citrix remote-access portal that had no multi-factor authentication, then deployed ransomware nine days later. Disconnecting the clearinghouse stalled pharmacy claims, medical claims and payments across the US.[1,2,42,43,44]
PathConnectivity & data → Human handoff
Air conditioning tripped at both trust data centres on the UK's record-heat day. Clinical IT went down and the trust ran on paper for weeks.[45,46]
A month-long ransomware attack on a health system with about 25% of regional inpatient discharges drove patients and ambulances to two unaffected academic EDs, raising their census, waits and stroke activations.[15,47]
PathConnectivity & data → Human handoff
Freezing weather knocked out generation and forced the largest controlled load shed in US history. Power loss spread to water systems and hospitals, and to patients at home on powered medical equipment.[48,49,12,11,14,13]
PathPower → Devices → Human handoff
A security incident led UHS to suspend user access to IT applications across its US operations; facilities ran on offline documentation for up to several weeks.[50,51]
Irma knocked out the transformer feeding a nursing home's air conditioning. 14 residents died; 12 deaths were ruled homicides.[52,53]
A self-spreading ransomware worm infected 34 English trusts and 603 primary-care and other NHS organisations, and at least 46 more trusts were disrupted. Thousands of appointments were cancelled and five hospitals diverted ambulances.[16,54,55]
PathConnectivity & data → Devices → Human handoff
Storm surge flooded basements holding fuel tanks and pumps at two Manhattan hospitals whose generators sat on upper floors. Both hospitals evacuated.[56,57,58]
After city power failed, Memorial ran on generators that failed as floodwater rose. 45 bodies were later recovered from the hospital.[59,60,61]
During the 2003 blackout multiple NYC hospital emergency generators failed. The outage was associated with about 90 excess deaths citywide.[62,63]
A network loop took down clinical applications at an academic medical centre for about four days, forcing a return to paper it had abandoned years earlier.[64,65,66,67]
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
| Instrument | What it requires |
|---|---|
| CMS Conditions of Participation, Emergency Preparedness, 42 CFR 482.15 | Hospitals must base their emergency plan on a facility-based and community-based all-hazards risk assessment, have arrangements with other hospitals to receive patients if operations are limited or stop, keep a communication plan with primary and alternate means, and run exercises at least twice a year.[17] |
| HIPAA Security Rule NPRM, 90 FR 898 (Jan 6, 2025), RIN 0945-AA22 (proposed, not final) | Proposes written procedures to restore critical systems and data within 72 hours, yearly written verification of business associates' technical safeguards, and business-associate notice within 24 hours of activating a contingency plan. It cites the Change Healthcare attack.[44] |
| NIST SP 800-161 Rev. 1 (May 2022, updated Nov 2024) | Guidance for building cybersecurity supply chain risk management into strategy, policy and risk assessment for the products and services an organisation buys, across organisation, mission and system levels.[69] |
| ASTP/ONC SAFER Guide: Contingency Planning (2025) | Self-assessment practices for EHR downtime, including paper forms for at least 8 hours, a tested read-only backup EHR, and unannounced downtime drills at least once a year. CMS requires hospitals to attest to the SAFER Guides annually.[67] |
| Joint Commission Sentinel Event Alert 67 (Aug 2023) | Guidance, not a standard. It calls on organisations to prepare all staff to keep care safe through an extended cyberattack downtime.[47] |
EU: the NIS2 Directive (EU) 2022/2555 requires essential and important entities to manage supply-chain security, including the quality and resilience of suppliers' products and services and cybersecurity terms in contracts with direct suppliers (Art. 21). It also provides for coordinated EU risk assessments of critical supply chains (Art. 22). UK: after Synnovis, the Cyber Security and Resilience (Network and Information Systems) Bill would let regulators designate 'critical suppliers' to essential services, and the government's factsheet uses Synnovis as its case study. The bill was still before Parliament in mid-2026. Netherlands: the Dutch Safety Board found in 2020 that hospitals' awareness of IT-failure risk had not kept pace with their dependence on IT. It recommended that hospitals map IT-to-care dependencies, test and drill regularly, and analyse serious outages in depth.[71,3,68]
FailSystems judgementJudgement, first draft. Likelihood 4: five large healthcare cascades in 2017–2024, three of them in 2024 alone, suggest the pattern recurs every year or two somewhere in the US or UK. Blast radius 5: cascades are by definition failures that spread beyond a single organisation; Change Healthcare and CrowdStrike reached national scale. Detectability 4: the triggering weakness is usually latent and often sits inside a supplier the hospital cannot see into, although once a cascade is running it is obvious.
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Cite this pageFailSystems. “Cascades.” https://failsystems.health201.com/layers/cascades/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.