AWS us-east-1 DynamoDB DNS failure disrupts cloud-hosted clinical systems
A race condition in DynamoDB's DNS automation broke a core AWS region for about 15 hours. Some cloud-hosted EHR users slowed or went to paper; others saw nothing.[22,24,25]
how automated healthcare fails, how you'd know, and what to do at each tier — every claim sourced, reviewed continuously
The networks, interfaces, vendors and records that carry clinical data, and what happens when data is missing, late or wrong.
This layer is everything between a clinician's question and the data that answers it: local networks and internet circuits, the EHR and its interfaces to lab, pharmacy and imaging, and the outside services a hospital depends on, such as claims clearinghouses and outsourced pathology. It also covers the data itself: whether it is present, current and correct.
It fails in two different ways. Data can be unavailable: a ransomware attack, a network loop or a vendor outage takes systems down, and staff know they are blind. Or data can be wrong: an order silently goes to a queue no one reads, a downtime copy is hours old, or results entered on paper never make it back. Unavailable data is loud and prompts a switch to backup processes; wrong data is quiet and does not.
Unplanned downtime is common. In one survey, 96% of large US health systems had at least one in three years, and 70% had one longer than 8 hours. Ransomware has made multi-week outages routine: about 44% of ransomware attacks on US care delivery organizations from 2016 to 2021 disrupted care, and in-hospital mortality rises among patients already admitted when an attack begins.
FailSystems viewIn a paper hospital, losing one department's records was a local problem. In an automated hospital, one identity system, one network core or one shared vendor carries every department's data, so failure is correlated and the backup is a mode of work nobody practises. We think the key distinction is 'unavailable' versus 'wrong'. Most planning targets the first: backups, warm sites, paper forms. The second defeats those plans because nothing tells anyone to use them. Defences against wrong data are reconciliation and monitoring (queues with owners, counts that must match, synthetic transactions), not redundancy.
Attackers encrypt servers and endpoints, often after days of undetected access and data theft. The organization then disconnects everything it cannot yet trust, so the EHR, lab, imaging, pharmacy and communications go dark together. Recovery is a rebuild, not a restart, and takes weeks.[1,2,3,4]
Warning signs
Seen inAscension ransomware and multi-week EHR downtime, Universal Health Services enterprise-wide IT shutdown, WannaCry ransomware across the NHS in England
A vendor that many organizations share (claims clearinghouse, pharmacy switch, outsourced pathology) is attacked or fails. Hospitals whose own systems are intact lose a function they cannot perform themselves, and every customer fails at once.[5,6,7,8]
Warning signs
Seen inChange Healthcare ransomware and national claims/pharmacy clearinghouse outage, Synnovis pathology ransomware, South-East London
The Joint Commission advises hospitals to be prepared to run with life- and safety-critical technology offline for four weeks or longer. Over days, order routing between departments, patient identification and result communication break down; lab turnaround slows and medication checks lapse. Back-entry after recovery creates a second risk period.[9,10,11,12,13]
Warning signs
Seen inAscension ransomware and multi-week EHR downtime, Synnovis pathology ransomware, South-East London
A loop, misconfiguration, carrier cut or failed core switch makes applications unreachable although servers and data are intact. Intermittent 'flapping' is worse than a clean outage because staff cannot tell whether to switch to paper.[14,15]
Warning signs
Orders, results or messages are accepted by one system and never reach the next, or land in a queue no one watches. The sender sees success, so no one switches to a backup process. Harm emerges as missed follow-up weeks later.[16,15,17]
Warning signs
Seen inVA Oracle Cerner EHR 'unknown queue' silently dropped clinical orders
Read-only downtime copies are snapshots and age from the moment the outage begins. After restoration, data captured on paper is back-entered late or not at all, and results produced during the outage may be absent from the electronic record. Clinicians decide on data that looks current but is not.[15,10,18]
Warning signs
Attackers target backup systems before encrypting, or backups turn out never to have been restored end to end. The organization then has no clean copy to restore from and must rebuild, or pay.[2,15,3]
Warning signs
Seen inChange Healthcare ransomware and national claims/pharmacy clearinghouse outage
When a system diverts ambulances and time-critical patients, nearby EDs absorb the load without extra staff. Waits, walk-outs and time-critical cases rise at hospitals that were never attacked; rural patients face much longer travel.[19,20,21]
Warning signs
Seen inWannaCry ransomware across the NHS in England, Universal Health Services enterprise-wide IT shutdown
A fault inside the provider (DNS automation, internal network congestion) disables core services across a region while the hospital's own building is fine. Impact depends on how each customer and each supplier built on the region: in October 2025 one Epic-on-AWS system slowed and another saw nothing, while NHS trusts using Oracle services went to paper.[22,23,24,25]
Warning signs
Seen inAWS us-east-1 DynamoDB DNS failure disrupts cloud-hosted clinical systems
A race condition in DynamoDB's DNS automation broke a core AWS region for about 15 hours. Some cloud-hosted EHR users slowed or went to paper; others saw nothing.[22,24,25]
A grid collapse cut power to continental Spain and Portugal for about ten hours. Hospitals largely held on generators; care outside them did not.[26,27,28,29,30]
A Class I software correction found that backlogged EHR-to-pump automated programming requests could load stale rate, dose or volume parameters.[31,32]
CISA and FDA reported that a low-cost patient monitor's firmware contained hidden functionality that could allow remote access and sent patient data to an external address; independent researchers later judged it an insecure design rather than an intentional backdoor.[33,34,35,36]
A faulty Rapid Response Content update to CrowdStrike's Falcon sensor crashed about 8.5 million Windows devices worldwide. Outside-in measurement found disrupted services at 759 of 2,232 US hospitals studied.[37,38,39,40,41,42,43]
PathDevices → Connectivity & data → Human handoff
Ransomware hit Synnovis, the pathology provider for several south-east London NHS trusts and GP practices. Blood testing and matching collapsed, more than 11,000 appointments and procedures were postponed, O-type blood ran short nationally, and one death was later partly attributed to a delayed result.[8,44,45,46,47,48,49,50]
PathConnectivity & data → Human handoff
A ransomware attack took Ascension's electronic records offline for about five weeks. Clinicians told KFF Health News of medication errors and delayed lab results, and one said he had no training for the attack; Ascension said its care teams were trained for such disruptions.[12,18]
Attackers used stolen credentials on a Change Healthcare Citrix remote-access portal that had no multi-factor authentication, then deployed ransomware nine days later. Disconnecting the clearinghouse stalled pharmacy claims, medical claims and payments across the US.[5,51,6,52,53]
PathConnectivity & data → Human handoff
Air conditioning tripped at both trust data centres on the UK's record-heat day. Clinical IT went down and the trust ran on paper for weeks.[54,55]
A month-long ransomware attack on a health system with about 25% of regional inpatient discharges drove patients and ambulances to two unaffected academic EDs, raising their census, waits and stroke activations.[19,13]
PathConnectivity & data → Human handoff
After go-live, the new EHR routed more than 11,000 clinical orders to a hidden queue instead of the intended service, without telling the ordering clinician; VHA identified 149 adverse events.[16]
A security incident led UHS to suspend user access to IT applications across its US operations; facilities ran on offline documentation for up to several weeks.[56,57]
A self-spreading ransomware worm infected 34 English trusts and 603 primary-care and other NHS organisations, and at least 46 more trusts were disrupted. Thousands of appointments were cancelled and five hospitals diverted ambulances.[58,21,59]
PathConnectivity & data → Devices → Human handoff
Ransomware made the EHR inaccessible; the hospital moved to paper within an hour and restored computers after 36 hours.[13,60,61]
A network loop took down clinical applications at an academic medical centre for about four days, forcing a return to paper it had abandoned years earlier.[14,62,63,15]
What should already be in place at each degradation tier for this layer. Tier 0 is normal automated running; tier 3 is paper, batteries and judgement.
These are practices reported or recommended in the cited sources, gathered for reference. They are not a prescription for your organisation; judge what fits your setting, and check the current official text of any standard.
| Instrument | What it requires |
|---|---|
| HIPAA Security Rule, 45 CFR 164.308(a)(7) Contingency plan | Covered entities must have a data backup plan, a disaster recovery plan and an emergency-mode operation plan; testing/revision and an applications-and-data criticality analysis are 'addressable'. A January 2025 NPRM would add written procedures to restore critical systems and data within 72 hours, but it was not final as of September 2026.[66] |
| CMS Hospital CoP Emergency preparedness, 42 CFR 482.15 | Hospitals must maintain a system of medical documentation that preserves patient information and keeps records available in an emergency, with the emergency plan and training/testing program reviewed at least every 2 years.[67] |
| ONC/ASTP SAFER Guide: Contingency Planning (2025 edition) | Self-assessment of 13 practices covering disaster recovery, generators, paper forms, tested backups, downtime training, independent communication, interface restart and downtime monitoring. CMS requires hospitals in the Medicare Promoting Interoperability Program to attest annually (yes or no) to completing all nine SAFER Guides.[15] |
| HHS 405(d) Health Industry Cybersecurity Practices (HICP), 2023 edition | Voluntary, sector-specific: ten practices against five threats including ransomware, scaled for small and large organizations in two technical volumes.[2] |
| HHS HPH Cybersecurity Performance Goals (CPGs) | Voluntary essential goals (e.g., MFA, incident planning, vendor cybersecurity requirements) and enhanced goals (e.g., network segmentation, third-party incident reporting, drilled incident plans).[7] |
| The Joint Commission Sentinel Event Alert 67 (2023) | Not a standard itself; recommends downtime planning committees, response teams, staff training and communication for extended cyber downtime, and points to TJC continuity-of-operations and disaster-recovery requirements.[13] |
In the EU the NIS2 Directive (2022/2555) keeps healthcare within its scope and imposes cybersecurity risk-management and incident-notification duties; ENISA's 2023 health threat landscape found ransomware in 54% of 215 reported health-sector incidents and a dedicated ransomware programme in only 27% of surveyed organisations. In England, DHSC's 2023-2030 cyber strategy aims for all health and social care organisations, including critical suppliers, to be cyber resilient by 2030. WannaCry (2017) and Synnovis (2024) are the reference cases: the first showed how unpatched systems and precautionary disconnection spread disruption, the second how a single pathology supplier can halt a region's diagnostics for months.[68,69,70,58,8]
FailSystems judgementJudgement: likelihood is 5 because unplanned EHR downtime is near-universal and ransomware attacks on care delivery roughly doubled between 2016 and 2021. Blast radius is 5 because shared vendors (Change Healthcare, Synnovis) and precautionary shutdowns take down whole regions or national functions at once. Detectability averages two extremes: outright outages are obvious (about 1), but silent misrouting and stale data can go unnoticed for months (about 5).
Each factor is scored 1–5 and multiplied, as in a classic FMEA risk priority number. This is our first-draft judgement, not a measurement; see how scoring works and how it will be revised.
These gaps drive what the nightly research pass looks for. If you have evidence, send it.
Cite this pageFailSystems. “Connectivity & data.” https://failsystems.health201.com/layers/connectivity/ (reviewed 2026-09-26). Health 201 / AstroNexus LLC. CC BY 4.0.
Information only, not advice. FailSystems is an aggregation and synthesis of published sources. It is not consulting, engineering, legal, regulatory or medical advice, and using it creates no professional relationship. Health systems are complex and no approach fits every organisation: anything you adopt is your own decision, at your own risk, and should be checked against the current official sources and by qualified people who know your setting. Full disclaimer.