Almost every healthcare business continuity plan is built to answer one question: what happens if we go down?
It’s the wrong question. Or rather, it’s only half of the right one.
A recent nationwide Health-ISAC business continuity exercise — a tabletop run across more than 500 participants from healthcare organizations around the country — posed a far harder version: what happens when a cloud identity provider that your organization shares with much of the rest of the sector goes down, and every hospital, payer, and vendor is scrambling for the same recovery resources at the same moment?
That scenario breaks most continuity plans, because most continuity plans are built in a silo. They model what happens inside your four walls and quietly assume the rest of the world keeps functioning normally while you recover. In a shared-vendor failure, that assumption is exactly backwards — and the gap between those two models is where healthcare organizations discover, too late, that their plan was never going to hold.
This dossier digs into the scenario the exercise surfaced, and how to build continuity planning that accounts for the environment you actually operate in.
The Scenario That Breaks the Plan
Picture the exercise scenario. A major cloud provider — the kind that hosts identity and credential services for organizations worldwide — begins experiencing performance degradation and a potential compromise of its credential service. Not a niche vendor. Think the scale of the hyperscale platforms that host identity for a huge share of the sector at once.
Here’s why that specific failure is so dangerous. Your credential service is your identity layer — it’s how everyone and everything logs in. And in a modern healthcare environment, your identity provider is the hub the entire technology stack connects through. Your applications authenticate against it. Your integrations depend on it. Your vendors and SaaS platforms tie into it. When the identity layer degrades, you don’t lose one system — you start losing everything, because everything downstream depends on being able to authenticate.
Now widen the lens, which is the entire point of the exercise. This isn’t happening only to you. It’s happening to every organization that shares that provider — which, for the major cloud platforms, means a substantial slice of the whole healthcare sector simultaneously. And it doesn’t stop at the organizations on the call. Their vendors and suppliers are probably on the same cloud infrastructure too. The failure trickles outward through every dependency, and the sector-wide blast radius is what a siloed plan never sees coming.
Why “How Long Can You Limp Along?” Is the Real Question
When the primary systems are down, most organizations fall back on manual workarounds. That’s the right instinct. But the exercise pressure-tested a question that rarely gets asked honestly: how long can the workarounds actually hold?
Manual processes are a bridge, not a destination. They work for a while, and then they degrade — staff fatigue, error rates climb, the volume outpaces what people can handle by hand. The critical planning question isn’t “do we have a manual workaround,” it’s “how many hours or days does that workaround remain viable before patient care is affected?” Most plans never put a number on it.
And the answer depends entirely on how critical your services are. This is where the exercise forced an uncomfortable clarity that every organization needs to internalize: not all healthcare organizations are equally critical, and in a sector-wide event, that ranking determines who gets resources first.
An organization with direct, immediate patient care — a hospital running critical care, an emergency department — sits at the top of that hierarchy. Patient safety is the deciding factor, every time. A health plan or payer, however important to the machinery of healthcare, does not have a patient bleeding in a bay. In a sector-wide resource crunch, hospitals take precedence. That’s not an opinion; it’s how triage works when the whole system is under stress. Every organization believes it’s critical. The exercise asks you to be honest about where you actually rank when the entire sector is contending for the same lifelines.
The Redirect Plan That Fails When Everyone Needs It
Here’s the assumption buried in most hospital continuity plans, and it’s the one the exercise detonates.
When a hospital’s systems go down and it can’t safely operate, it has a plan: stabilize, stand up emergency capacity, and if necessary, redirect patients to another hospital. Most hospitals know how to do this. It’s a standard piece of continuity planning.
But that plan contains a silent dependency: the other hospital has to be up.
In a sector-wide event — the shared cloud provider, the widespread ransomware campaign — the hospital you planned to divert to is fighting the same fire you are. So is the one after that. The redirect plan assumed a functioning neighbor, and in a systemic failure there isn’t one. Where does the patient go? That question has no good answer if your plan stopped at “divert,” and it’s precisely the question a siloed tabletop never raises.
The Incident-Response Queue Nobody Plans For
This is the detail from the exercise that should change how every healthcare organization thinks about its third-party incident-response arrangements.
Many organizations carry an incident-response retainer with a third-party firm — a contract guaranteeing expert help when something goes wrong. It feels like insurance. On a normal day, it is.
But run the math the exercise ran. Suppose your IR firm has ten analysts. On an ordinary day, when your incident is one of a handful they’re handling, that retainer gets you fast, dedicated help. Now suppose a shared cloud provider goes down and a hundred of that firm’s client organizations are all breached or degraded at the same moment. Ten analysts. A hundred simultaneous emergencies. Where are you in that queue?
The retainer doesn’t move you to the front. It guarantees you access to a resource pool that just became catastrophically oversubscribed. In a sector-wide event, the very thing that makes the event sector-wide — everyone using the same handful of vendors — means everyone is calling the same handful of responders at once. The resource you were counting on is the resource everyone else was counting on too.
This is concentration risk applied to recovery, and almost no siloed continuity plan models it. The plan assumes your retainer is available. It doesn’t ask whether it’s available when a hundred other victims have the same retainer with the same firm.
// INCOMING TRANSMISSION
Status: Secure Episode 026 — World Leaks Extortion, Quishing Attacks, and Inside the Latest Health-ISAC Threat Briefing takes you inside the nationwide Health-ISAC business continuity exercise firsthand. Our CISO was in the room for the shared-cloud-provider scenario and walks through what 500+ healthcare organizations learned about limping along, resource queues, and why patient safety sets the recovery priority. Listen for the operator's view on defending as a sector, not alone.
INITIATE PLAYBACK »Why Siloed Continuity Planning Is the Root Failure
Step back and the pattern is clear. Every problem above — the limp-along ceiling, the failed redirect, the oversubscribed IR retainer — comes from the same root cause: most organizations run their business continuity exercises in isolation.
They model what happens inside their own organization. They don’t model the environment they operate in. They plan as if their failure is the only failure, when the most likely large-scale disaster is precisely one that hits many organizations at once through a shared dependency — a cloud provider, a widely used software platform, a common managed service.
The Health-ISAC exercise exposed that blind spot by design. It wasn’t trying to prepare one organization. It was trying to prepare the sector — because the threats that matter most now are the ones that don’t hit one victim at a time. This is why sector-level intelligence sharing and sector-level exercises exist: no single organization can see the whole board, and the failures that will hurt most are the shared ones.
How to Build Continuity Planning That Accounts for the Sector
Here is how to close the gap the exercise revealed.
Vary who runs your tabletop.
If the same internal team runs your business continuity exercise every year, you rehearse the same assumptions every year. You cannot predict how a real event will unfold, so the value of an exercise is in being surprised — in having someone introduce a scenario your team would never have thought to model. Bring in an outside facilitator periodically, precisely because they’ll run it differently, inject scenarios you’re blind to, and pressure-test the assumptions your team has stopped questioning. Different facilitators bring different disasters. That’s the point.
Model the environment, not just the org.
Explicitly build sector-wide scenarios into your planning. What happens if your cloud identity provider degrades? If a software platform used across healthcare is compromised? If your IR firm’s other clients are all down at once? These are not exotic edge cases — they are the shape of the most consequential incidents now. A plan that only models your own isolated failure is planning for the wrong disaster.
Put a number on your limp-along window.
For each critical service, determine honestly how long your manual workaround remains viable before patient care is affected. That number drives everything — your recovery-time priorities, your resource decisions, your escalation triggers. “We have a manual process” is not a plan. “Our manual process holds for roughly 12 hours before the ED backs up” is.
Understand where you rank.
Map your services by criticality against the sector, not just internally. Know which of your functions provide direct patient care and which don’t, because in a resource-constrained sector-wide event, that ranking determines your realistic access to shared resources. Plan for the position you’ll actually be in, not the priority you wish you had.
Diversify — or at least map — your concentration risk.
You may not be able to avoid the hyperscale cloud providers; few can. But you can know your single points of failure, understand which vendors you share with the rest of the sector, and build contingency for the specific case where a shared vendor takes down you and your neighbors simultaneously. Even where you can’t eliminate the concentration, naming it changes how you plan around it.
Get eyes on the ground through threat intelligence.
The exercise itself was only possible because these organizations were plugged into Health-ISAC — the sector’s shared intelligence network. You need eyes watching the environment for you, because you cannot watch all of it yourself. Sector intelligence is how you learn a shared vendor is degrading before it becomes your incident. (For the full breakdown of the intelligence-sharing model and the threats the sector is tracking right now, that’s the subject of the companion episode above.)
The Staffing Reality Underneath All of It
One hard truth threaded through the entire exercise: none of this works without people.
The recurring root cause of healthcare security failure isn’t technology — it’s staffing and budget. Intelligence you can’t act on because you’re understaffed is intelligence wasted. A continuity plan nobody has the capacity to execute is a document, not a defense. And smaller and rural organizations feel this most acutely — many increasingly lean on external threat-intelligence networks and outside partners precisely because they cannot build these capabilities in-house, and realistically never will.
That’s not a failure. It’s a rational response to an asymmetry. No individual healthcare organization can out-resource organized cybercrime and nation-state actors alone. The organizations that fare best are the ones that recognize this early and build the external relationships — intelligence networks, exercise partners, response capacity — before the event, not during it.
Marching Orders
Run a sector-aware tabletop this quarter.
Model a shared-vendor, sector-wide failure — a cloud identity provider degradation is the ideal scenario. Bring in an outside facilitator to run it so you’re tested on assumptions your own team can’t see.
Quantify your limp-along window for every critical service.
Turn “we have a manual workaround” into a specific number of hours, and let that number drive your recovery priorities.
Pressure-test your IR retainer against concentration risk.
Ask your incident-response firm directly: what happens to your response time if a large share of your clients are hit simultaneously? Plan for that answer, not the best-case one.
Map your single points of failure and shared vendors.
Know exactly which dependencies you share with the rest of the sector, and build specific contingency for their simultaneous failure.
Execute the Standard
The most dangerous disasters in healthcare are no longer the ones that hit you alone. They’re the ones that hit the whole sector through a shared dependency — and those are exactly the disasters that siloed business continuity planning is structurally blind to. The redirect that fails because the next hospital is down. The retainer that’s worthless because a hundred other victims hold the same one. The manual workaround with no honest expiration date.
You cannot predict how the next event will unfold. You can only prepare for a wider range of them than your competitors do — and you can refuse to plan as though you operate alone, because you don’t.
If your organization needs to build a business continuity plan that models sector-wide and shared-vendor failure, run a tabletop exercise with an outside perspective that tests the assumptions your team can’t see, or map the concentration risk hiding in your vendor stack, that is the work we do. Verify your security posture at watchur6.com/secure, or establish a secure line at watchur6.com/contact.
Trust but verify your own posture. Model the sector, not just yourself. Put a number on your limp-along window. Know where you rank. Execute the standard.
This Sitrep draws on themes from a nationwide Health-ISAC healthcare business continuity exercise. Health-ISAC (Health Information Sharing and Analysis Center) is the healthcare sector’s member-driven threat-intelligence sharing organization. Specific exercise content shared under member restrictions is not reproduced here.