Resiliency
Resiliency is the word institutions reach for when they mean toughness — redundancy, backups, the capacity to take a hit and carry on. That is part of it, and it is the smaller part. The definition I work from is narrower and more demanding: resiliency is less about preventing failure than knowing, in advance and in detail, the shape failure will take. This is a short map of what that means and how the discipline is practiced.
The wrong question
Prevention asks: how do we keep this from happening? It is a necessary question with a known ceiling. Complex, tightly coupled systems fail in ways no component analysis predicts; adversaries adapt faster than controls; and the failure that finally gets you is, almost by definition, the one your defenses were not built against. Resiliency asks the question after that one: when this fails — and it will — what exactly happens next? Who notices, through what signal, how fast. What stops working, what keeps working, what must never stop. What will we need in hand that cannot be improvised on the day. Prevention is a wager that you will be smarter than your failure. Resiliency is the humbler position: assume the breach, the outage, the bad morning — and win on the details.
Failure has a shape
The word detail is doing the work. Every institution “knows” it could be breached or go down; that knowledge is cheap and changes nothing. What changes things is specificity: the dependency nobody mapped, the escalation path that dead-ends at a voicemail, the recovery tool that requires the very network that is gone, the credential rotation that takes down half your own services when you finally run it. Failures are not abstractions. They have anatomy — sequence, timing, dependencies, second-order effects — and the institutions that come through are the ones that studied the anatomy in advance. The generic anticipation of failure produces documents. The detailed anticipation of failure produces capabilities. I wrote about a vivid recent demonstration — the first fully autonomous cyber attack, and the pre-built capabilities that saved its victim — in The shape of failure.
The craft
Foreknowledge of failure sounds like a gift. It is a product, and three instruments manufacture it.
Exercises. A plan is a stack of assumptions that has never been contradicted. Rehearsal is the instrument that contradicts them early, while it is still cheap — that lets you watch the handoff stall and the “isolated” system keep talking at exercise prices rather than production prices. The point of a crisis exercise is not to practice the plan. It is to meet the failure before it has cost you anything.
Assessments. A system diagram shows how a thing is meant to behave. An assessment reads it the other way — as a set of assumptions nobody has tested, each one a place where the failure will differ from the plan. The discipline is asking, of every control and dependency: what would have to be true for this to hold, and who has checked.
After-action reviews. Every incident and every exercise ends with a choice: harvest the specifics honestly, or file a summary that flatters the plan. The after-action review is where imagined failure gets replaced by observed failure — and an institution that skips it pays for the same lesson twice.
What holds up
- Assume failure. Design from the failed state backward. The question is never whether the control can fail but what the morning after looks like — and whether you have already met it.
- Detail over documents. A binder describes a failure in general. Capability exists only where someone has worked out the specific one — and tested the response against it.
- Degrade gracefully. Nothing important should be all-or-nothing. A resilient system loses capability in a planned order, keeping its essential services alive while the rest waits.
- Rehearse the people. Tooling recovers systems; people recover institutions. Decision rights, thresholds, and working relationships have to exist before the crisis — it is a poor moment to be exchanging introductions.
None of this is pessimism. Assuming failure is not fatalism; it is the only posture from which the details can be gotten right. And foreknowledge, unlike prevention, degrades gracefully itself: even a partially imagined failure meets a partially prepared response. Prevention buys you time. Resiliency is what you have when the time runs out.
Further reading
- Charles Perrow, Normal Accidents — why complex, tightly coupled systems fail in ways no component analysis predicts.
- Erik Hollnagel, David D. Woods & Nancy Leveson (eds.), Resilience Engineering: Concepts and Precepts — the field's founding argument: safety is something a system does, not something it has.
- Sidney Dekker, Drift into Failure — how ordinary, locally sensible decisions accumulate into systemic failure.