Skip to content

Your DR plan has never been tested. Here’s what that costs.

An untested recovery plan is a document, not a capability. The gap between the two is discovered at the worst possible moment, and it is measurable in advance.

6 min read Continuity

Ask an IT director whether the organisation has a disaster recovery plan and the answer is almost always yes. Ask when it was last invoked, in anger or in rehearsal, and the room goes quiet. In our experience the plan exists, the secondary site is paid for monthly, and nobody has ever failed over to it.

That is not negligence. Testing recovery means deliberately breaking something that is currently working, usually outside business hours, with the person who authorised it carrying the risk personally. The incentives run against it every quarter until the quarter when they do not.

The three numbers that are usually wrong

Recovery objectives get set early in a continuity project and are rarely revisited. Three numbers deserve scrutiny, because untested plans get all three wrong in the same direction.

  • Recovery time objective — how long the business says it can be down. Usually set by ambition rather than by architecture.
  • Recovery point objective — how much data the business can afford to lose. Usually stated as zero, and almost never delivered as zero.
  • Actual recovery time — what the restore genuinely takes, including the parts nobody counted.

The third is the one that surprises people. A four-hour RTO on paper becomes eleven hours in a rehearsal because the plan measured the database restore and forgot DNS propagation, the certificate that lives on the primary load balancer, the batch job that must be rerun in sequence, and the forty minutes spent finding out who has the vault password.

What the cost actually looks like

Boards respond to a number, so give them one. Take the revenue-generating systems, work out the transaction value per hour, and multiply by the difference between the stated RTO and the real one. For a mid-sized payments business we worked with, the stated objective was two hours; the first honest rehearsal took nine. Seven hours of unplanned exposure, repeated across a realistic outage frequency, made the case for investment in a single slide.

Add the costs that do not appear on the revenue line: regulatory notification where availability is a supervisory concern, contractual service credits, the staff cost of a manual workaround, and the reputational cost of being the institution that was down when a competitor was not.

Start with a tabletop, not a failover

The objection to testing is risk, and it is a fair objection. The answer is to sequence the rehearsals so that each one is cheap enough to authorise.

  1. Tabletop walkthrough. Everyone in a room, one scenario, no systems touched. Two hours. It will find missing contact details, unclear authority to declare, and steps that depend on one person.
  2. Component restore. Restore one non-critical database to a scratch environment and time it honestly, including the parts before and after the restore itself.
  3. Isolated failover. Fail one service over to the secondary site during a maintenance window, with a rehearsed rollback.
  4. Full failover with the business present. Only once the first three run clean. This is the one that produces the evidence a regulator or auditor will accept.

Most of the value arrives at step one. The tabletop is the cheapest control test available in any discipline, and it is the one most consistently skipped.

What a rehearsal should produce

A test that produces only a pass is a test that was designed to pass. A useful rehearsal produces a list of things that did not work, each with an owner and a date, and a revised recovery time based on what actually happened rather than what was planned.

It should also produce a corrected plan. The version that goes back in the drawer must be the one edited during the exercise, with the steps in the order they were really performed, written plainly enough to follow at three in the morning by whoever is on call.

The failure modes we see repeatedly

  • Backups that complete successfully and cannot be restored, because nobody has ever restored one.
  • A secondary site running a configuration that drifted from primary eighteen months ago.
  • Recovery documentation stored on the system being recovered.
  • Credentials held by one administrator who is, by definition, on leave.
  • A supplier whose contracted response time was never tested and turns out to start when a ticket is triaged, not when it is raised.
  • Staff who know the plan exists and have never read it.

Each of these is found by a rehearsal costing a fraction of the outage it prevents. None of them is found by reviewing the document.

A plan nobody has rehearsed is an assumption with a version number.

If your organisation has not tested recovery in the last twelve months, the honest position to take to the board is that recovery time is currently unknown. That statement is uncomfortable, and it is more useful than a number nobody has verified.

Published 21 July 2026. This is general commentary, not advice on your circumstances, and regulatory positions move. Check anything here against your own obligations before acting on it.

Follow-up

Bring us the version of this that is on your desk.

If this briefing describes a problem you currently have, a short call is more useful than another article.

Send an enquiry

+254 721 687846 taarifa@techrisk.co.ke