Updated: October 2026
Your backup dashboard has been green for months. That tells you the data got copied, and nothing about whether you can run on it by lunchtime. Here's how to set up disaster recovery testing that answers that: which tests to run, how often, what counts as a pass, and what to write down.
TL;DR
- Definition. Disaster recovery testing proves you can restore systems within your RTO and RPO.
- Five types. Checklist review, tabletop, walkthrough, parallel test and full failover, from cheapest to most disruptive.
- Schedule. Automated checks nightly, a real restore monthly, a timed boot test quarterly, a failover yearly.
- Pass. Restore time under the RTO, restore point inside the RPO, and a user who can log in and work.
- Proof. Every test ends in a one-page report with a named owner for each fix.
What Is Disaster Recovery Testing?
Disaster recovery testing is the practice of restoring systems from backup or failing them over to a secondary site, on purpose, and timing how long it takes. A test passes when the business can use the restored systems within its recovery time objective (RTO) and without losing more data than its recovery point objective (RPO) allows.
That's a different question from a backup check. A backup check asks whether the data exists. A DR test asks whether the business can run again, which pulls in everything the data depends on: domain controllers, DNS, licensing, network paths, and the people who know the steps.
It also sits inside a bigger plan. The DR plan says how you get systems back. The business continuity plan says how people keep working in the meantime, and the incident response plan says how you contain the attack first. If those three blur together, our DR plan vs IR plan breakdown separates them.
Why a Green Backup Job Proves Less Than You Think
Confidence and results sit a long way apart. Veeam's 2026 ransomware report found 90% of organizations were confident they could recover. Among ransomware victims, only 28% got all of their affected data back.
The testing gap shows up in the tooling too. Acronis looked at its own platform telemetry for Q1 2026 and found that 82% of backup rules had no automated test schedule. Only 18% ran monthly tests. That's rules on one vendor's platform, not organizations, but it's a useful mirror.
Even the big clouds get caught by this. When Azure Front Door went down on October 29, 2025, Microsoft's own write-up says the bad configuration also updated the "last known good" snapshot. The rollback target was broken too. Nobody finds that in a backup report. You find it when you try to roll back.
This r/sysadmin thread asks the question every green dashboard hides: how do you test the restore, not the backup? The replies are a good tour of what people check and what slips through.
The 5 Types of Disaster Recovery Tests
These disaster recovery testing methods run from a desk exercise to switching production off. Each one proves more than the last, and costs more.
- Checklist review. The team reads the DR plan and checks it against reality: contacts, systems, passwords, vendor numbers.
- Tabletop exercise. The team talks through a scenario, step by step, and finds the gaps in who does what.
- Walkthrough (simulation). Techs perform the recovery steps in an isolated environment, without touching production.
- Parallel test. Systems are recovered at the DR site and run alongside production, which keeps serving users.
- Full failover (interruption) test. Production is switched off and the business runs on the recovered systems.
NIST uses different labels. SP 800-34 talks about tabletop, functional and full-scale exercises, but the ladder is the same.
Picking one is a trade between what it proves and what it risks. Here's how they compare.
| Test type | Effort | Disruption risk | What it proves | How often |
|---|---|---|---|---|
| Checklist review | Low | None | The plan is current | Quarterly, and after any staff or vendor change |
| Tabletop exercise | Low to medium | None | People know their roles | Twice a year |
| Walkthrough | Medium | Low (isolated network) | The steps work and are timed | Quarterly for critical systems |
| Parallel test | High | Low to medium | The DR site can carry the load | Yearly for critical systems |
| Full failover | Very high | High | The business runs on DR | Yearly or less, for one system group at a time |
The frequencies are a starting point, not a standard. Tie them to how much downtime each system can take, which is the next section.
W. Curtis Preston, who's written about backup and recovery for decades, walks through the same ladder here, from a single file restore up to full DR, and how to test without putting production at risk.
A Testing Schedule by System Tier
Not every system needs the same schedule. Sort them by RTO first. Then give each tier layers that stack: cheap automated checks often, expensive human tests rarely.
| Layer | Tier 1 (RTO 1 hour) | Tier 2 (RTO 4 hours) | Tier 3 (RTO 24 hours) |
|---|---|---|---|
| Automated verification (boot, heartbeat) | Nightly | Nightly | Weekly |
| Manual restore of a file, mailbox or database | Monthly | Monthly | Quarterly |
| Timed boot test against the RTO | Monthly | Quarterly | Twice a year |
| Walkthrough or tabletop | Quarterly | Twice a year | Yearly |
| Parallel or full failover | Yearly | Yearly, as part of a group | Only if the business asks |
The automated layers do the repetitive checking. That frees tech hours for the tests that need a human, like timing a boot or talking through who calls the insurer.
The monthly restore is where rotation helps. In this r/sysadmin thread, the top answer describes a script that picks a random team member each month and opens a ticket for four restores: a VM, a single file, a storage volume and a test failover. Nobody gets to be the only person who knows how.
If you're still setting RTOs and backup scope, our BCDR guide covers that groundwork first.
How to Run a Disaster Recovery Test, Step by Step
A DR test plan doesn't need to be long. It needs to be written down before anyone touches a console.
- Scope it. Name the systems, the scenario and the test type.
- Set the targets. Write the RTO and RPO for each system in the plan, so the pass mark exists before the test does.
- Prepare the rollback. Isolate the test network, and know how you'll undo it if something leaks into production.
- Run it and time it. Start the clock at the moment of the simulated failure, not when the restore job begins.
- Validate at the application level. Have a real user log in and complete a normal task.
- Record everything. Times, errors, workarounds and every step the runbook got wrong.
- Fix and retest. Assign each problem an owner and a date, then run the failed part again.
Time your first test carefully. There's no reliable benchmark for how long a DR test takes, so your own first run becomes the baseline you plan against.
What Counts as a Pass
"It restored" isn't a pass. A pass has to be defined before the test starts, or every test passes. AWS's Well-Architected guidance frames it well: the RTO and RPO are met when the workload is back to the specified state in the specified time.
A useful pass has five parts. The measured restore time is under the RTO, counted from the simulated failure. The restore point is newer than the RPO allows. A user can log in and do their job, not just see a login screen. The dependencies came back: DNS, domain controllers, license servers, certificates, identity and MFA. And someone other than the runbook's author can follow it, which is exactly how NIST says plans should be written: for "personnel unfamiliar with the plan."
There's one more check for ransomware scenarios. Mandiant's M-Trends 2026 puts the global median dwell time at 14 days. A restore point from last Tuesday may already contain the attacker. Practice finding a clean point, not just the newest one.
Tests That Pass When They Shouldn't
Automated verification is useful. It also proves less than its green tick suggests, and the documentation says so if you read it.
Datto's screenshot verification boots the backup and captures the login screen. It does that with no network adapter attached, so it proves the OS boots, not that the apps talk to each other. Veeam's SureBackup is more thorough, but its predefined domain controller test checks that port 389 answers. A DC can answer on 389 and still be broken.
Hardware is another trap. A bare-metal restore to different hardware without injected drivers comes up with default Windows drivers only. That's fine until the storage or network controller needs something else.
Then there's scope drift. A VM gets built in March and never added to the backup job. The test passes because it tests what's in the job. Reconcile the backup scope against your live asset list every month. OpenFrame's device inventory, pulled from the endpoints through osquery, is one way to get that list without a spreadsheet. The backup tool itself stays whatever you run today.
The last false pass is human. A test run by the person who wrote the runbook will pass because they fill the gaps from memory. Rotate who runs it.
Disaster Recovery Testing Checklist
Use this as the skeleton for every test. It covers the before, during and after, and it doubles as your DR test plan template.
| Phase | Check | Done |
|---|---|---|
| Before | Scope, scenario and test type written down | ☐ |
| Before | RTO and RPO recorded for every system in scope | ☐ |
| Before | Backup scope reconciled against the current asset list | ☐ |
| Before | Isolated network ready, rollback plan agreed | ☐ |
| Before | Business owner told, change window booked | ☐ |
| During | Clock started at the simulated failure | ☐ |
| During | Restore point chosen and its age recorded | ☐ |
| During | Dependencies restored in order: identity, DNS, databases, apps | ☐ |
| During | Every deviation from the runbook noted | ☐ |
| After | A user logged in and completed a real task | ☐ |
| After | Measured time and data loss compared to targets | ☐ |
| After | Test environment torn down, nothing left running | ☐ |
| After | Report written, fixes assigned, retest date set | ☐ |
Scenarios Worth Testing
A test is only as useful as the disaster it rehearses. Rotate through these.
Ransomware. Assume the newest restore points are infected and the backup console is a target. Marks & Spencer paused online orders on April 25, 2025 and resumed them 46 days later. Jaguar Land Rover stopped production around September 1, 2025 and began a phased restart on October 8.
A deleted mailbox or SharePoint site. Retention settings aren't backup, and Microsoft's own guidance treats them separately. Test a single-mailbox restore monthly. Our Microsoft 365 backup guide covers what's protected by default and what isn't.
A dead host or storage array. The classic. Time a full VM restore to different hardware, drivers included.
A cloud region outage. AWS us-east-1 was degraded for about 14.5 hours on October 19 and 20, 2025. If a Tier 1 system lives in one region, test what happens when that region doesn't answer.
Losing the backup tool itself. The backup server is encrypted, or the vendor has an outage. Can you restore from the offsite copy with nothing but the documentation?
The Disaster Recovery Test Report
A test that isn't written down didn't happen, at least as far as an auditor or your leadership is concerned. The format doesn't need inventing. NIST's SP 800-84 and FEMA's HSEEP after-action report both use the same bones: objectives, what happened, what went wrong, and an improvement plan with owners and dates.
| Field | Example (illustrative) |
|---|---|
| Date and test type | 14 Oct 2026, walkthrough |
| Systems in scope | File server, ERP database, domain controller |
| RTO target vs measured | 60 min target, 41 min measured |
| RPO target vs restore point age | 4 hours target, 2 hours 10 minutes |
| Application check | Finance user posted an invoice in ERP |
| Result | Pass |
| Issues found | Runbook missing the license server restart step |
| Fix, owner, due date | Update runbook, J. Rivera, 21 Oct 2026 |
| Next test | January 2027 |
For leadership, compress it to one line per Tier 1 system: target, measured, pass or fail. That line is the whole point of the exercise.
Who Expects Proof of Testing
Auditors and insurers ask for the report. HIPAA's Security Rule, at 45 CFR 164.308(a)(7)(ii)(D), calls for "procedures for periodic testing and revision of contingency plans." It's an addressable specification today. A 2025 proposed rule would require testing at least every 12 months, but it isn't final yet.
SOC 2's availability criterion A1.3 expects the organization to test its recovery plan procedures. PCI DSS 12.10.2 requires the incident response plan, including backup processes, to be tested at least once every 12 months.
Insurers ask too. Beazley's ransomware supplemental application asks, in question 43, "How frequently do you perform a test restoration from backups?" The answers range from "Never/not regularly" to "Quarterly or more often." Your test reports are what backs up the box you tick.
Start With One System
Disaster recovery testing comes down to a schedule, a stopwatch and a report someone reads. The automated layers catch broken backups early, and the human layers prove the business can run.
Start small this week: pick one Tier 1 system, run a timed walkthrough, and write the one-line result. For the tools side of recovery, our roundup of MSP backup solutions is the next read.
Conrad Lunderstedt
Solution Architect
I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.
