Back to Blog
    Downtime Procedures Are Useless If They Live in the System That's Down
    problem
    IT director

    Downtime Procedures Are Useless If They Live in the System That's Down

    Nobody designed it this way. The instructions ended up inside the thing the instructions are about, because putting documents in the sanctioned place is what a conscientious person does.

    Solution Compass
    June 30, 20265 min read

    Ask your informatics team where the EMR downtime procedure is documented. There is a good chance the honest answer involves a link on the intranet.

    Now ask what the intranet runs behind. In most health systems it is the same directory, the same sign-on, and frequently the same virtual infrastructure as the thing that is down.

    I have watched this go wrong twice, and both times it went wrong in the same undramatic way: nobody could not find the procedure, exactly. They found it eventually. It took eleven minutes, during which the unit did the thing they remembered from the last downtime, which was mostly right.

    The dependency nobody drew

    Business continuity planning in healthcare is generally taken seriously. There are plans. They are reviewed. Somebody owns them.

    What is rarely done is drawing the dependency graph for the plan itself, as distinct from the systems the plan covers. The plan says: when the EMR is unavailable, use the downtime forms and follow the documented workflow. Fine. Where are the downtime forms? A shared drive. What authenticates the shared drive? The same directory service. What happens to that directory service in the failure modes you are actually planning for?

    For a straightforward application outage, nothing — the EMR is down and everything else is fine, and this whole concern is theoretical. That is the scenario everyone tests, because it is the easiest to simulate and the least disruptive to arrange.

    For an infrastructure event, a network segmentation failure, or the ransomware scenario that everyone is now obligated to plan for, the answer is different, and the difference is the entire point. In those cases the documentation and the system it documents fail together, because somebody stored the instructions inside the thing the instructions are about.

    Downtime Procedures Are Useless If They Live in the System That's Down

    Nobody designed it that way. It happened because the intranet was the sanctioned place to put documents, and putting documents in the sanctioned place is what a conscientious person does.

    Printed copies are not the answer people think

    The traditional mitigation is a binder. Every unit has one. It is on a shelf, it has a cover sheet with a revision date, and it satisfies the requirement.

    I am not against binders and I would rather have one than not. But I want to name honestly what they do and do not solve, because the binder is frequently treated as closing the issue when it does not.

    A binder solves reachability. It does not solve currency, and the gap between those two grows continuously and invisibly. The binder was assembled at a point in time. Since then the EMR has been upgraded twice, a form has changed, a phone number belongs to somebody else, and a workflow was revised after an event. None of that reached the shelf, because updating the binder requires someone to print pages, walk to every unit, and swap them, and that job belongs to nobody.

    There is a second problem, which is that the binder is organized as a document rather than as an answer. It is forty pages with a table of contents. The person opening it needs one specific thing — how to handle a stat lab order during downtime — and they are paging through a plan written for the person who wrote the plan.

    So the binder is a reasonable floor and a poor ceiling. Keep it. Do not mistake it for a solution.

    Design for the failure you are actually planning for

    The more useful exercise is to be explicit about which failure modes you are covering, and to accept that different ones need different answers.

    For a single application outage, an intranet page is fine, and the honest thing to do is say so and stop over-engineering it.

    For an infrastructure or identity event, the requirement is that the content is reachable without your network and without your directory. That is a genuinely different requirement and it usually means content held somewhere outside your primary environment, reachable from a phone on cellular, with an authentication path that does not depend on the system that is down. Which is uncomfortable, because the reflex in a security review is to demand exactly the coupling that makes it fail.

    Downtime Procedures Are Useless If They Live in the System That's Down

    That tension is real and I do not want to pretend it resolves neatly. Content reachable when your directory is down is, by construction, content protected by something other than your directory. The way through it is scope: this is operational procedure, not patient data, and the material risk of a downtime workflow being read by the wrong person is very low compared to the risk of it being unreachable by the right one. Make that argument explicitly and get it decided, rather than letting the default answer settle it silently.

    And whatever you build, someone has to be able to find the specific answer inside it under pressure. Retrieval matters more here than anywhere else in the building — the person searching is stressed, working around a broken system, and typing whatever comes to mind. That is the case Solution Compass is built for, and the reason offline reachability is a design question rather than a feature to be added later.

    Test retrievability, not just the plan

    Tabletop exercises test decisions. Someone reads a scenario, the group discusses, actions are assigned, the exercise is documented. That is worth doing and it is not what I am asking for.

    Add one step. Partway through, pick a person who is not in the planning group — a night charge nurse, a newer analyst — and ask them to produce the specific document that covers the specific situation. Give them the tools they would actually have in that scenario, which may mean their phone and no network. Time it.

    You will learn three things in about four minutes. Whether the document is reachable. Whether it is current. And whether the person can find the relevant part of it, which is a separate question from whether it exists and is the one that usually fails.

    Run it twice a year and vary who you pick. If you always pick the person who helped write the plan, you are measuring their memory rather than your documentation.

    The eleven minutes I mentioned earlier were not spent searching. They were spent by two people who each believed the other knew where it was, deferring to each other, while a third quietly started doing it from memory and got most of it right. That is the actual failure mode, and no plan document has a section for it.

    Ready to stop losing institutional knowledge?

    Solution Compass captures your organization's collective wisdom and makes it instantly searchable.