Updated: October 2026
When a server room floods, the backup proves the data survived, but the servers still need somewhere to run. Getting them running again is the slow and expensive part, and a backup on its own doesn't solve it. Here's how disaster recovery as a service (DRaaS) works, what it costs, and how to tell whether you need it.
TL;DR
- DRaaS keeps a running copy of your servers somewhere else. It replicates machines to a provider's cloud and starts them there when your site is down, then moves them back later.
- It sits between backup and a second data center. Backup gets data back in hours or days. DRaaS gets servers running in minutes to hours.
- Replication frequency sets your RPO. Everything else sets your RTO. Boot order, networking, DNS and people decide how fast users are back, not the replication engine.
- The per-server fee is under a third of the bill. Replica storage, snapshots, replication servers and compute during tests and failover make up the rest.
- Test failovers are the product. Budget for them in the contract and run them against the RTO you promised.
What Is Disaster Recovery as a Service?
Disaster recovery as a service is a copy of your servers that lives at a provider and can be switched on when your own site can't run them. The provider keeps that copy current, supplies the compute to run it, and gives you a way to move back once the original site recovers.
Four parts make it work. Replication copies changes from your servers to the provider, usually at the block level and close to continuously. A recovery environment holds the replicas on cheap storage with minimal compute until you need them. Orchestration starts the machines in the right order, with the right networks and addresses. Failback moves everything home when the emergency is over.
Microsoft describes its own service, Azure Site Recovery, in exactly those terms: it "replicates workloads running on physical and virtual machines (VMs) from a primary site to a secondary location," you fail over when the primary site is down, and you "fail back to it" once it's running again. AWS Elastic Disaster Recovery works the same way, with a low-cost staging area that keeps replication running and full recovery instances launched only for drills or a real recovery.
The "as a service" part is the point. You don't buy, rack and patch a second set of servers that sit idle for years. You rent the recovery site and pay mostly when you use it.
DRaaS vs Backup vs BaaS vs BCDR
These four terms get used as if they were interchangeable. They answer different questions.
| What it protects | Where you recover | Typical recovery time | |
|---|---|---|---|
| Backup | Copies of data and system images | Hardware you still have or can buy | Hours to days |
| Backup as a service (BaaS) | The same copies, stored and managed by a provider | Still your own hardware, or a restore into cloud storage | Hours to days |
| DRaaS | Replicated servers, ready to boot | The provider's cloud, on their compute | Minutes to hours |
| BCDR | The whole plan: people, processes, tools | Wherever the plan says | Whatever the plan commits to |
Backup and DRaaS are not alternatives. Replication faithfully copies whatever happens on the source, including a mistaken deletion or a ransomware encryption. You still need point-in-time backups, ideally immutable, to go back further than the replication history reaches.
BCDR is the program that decides which systems need which kind of protection. DRaaS is one of the tools inside it, usually reserved for the systems the business can't be without for more than a few hours. Our BCDR guide covers how to build the plan around it.
The cost of getting this wrong is not abstract. In the Uptime Institute's 2025 survey, published in its 2026 outage analysis, 57% of respondents said their most recent major outage cost more than $100,000, and one in five put it above $1 million.
This IBM Technology explainer walks through the backup and disaster recovery split in under ten minutes:
How Failover and Failback Work
A failover is the moment you stop running on your own site and start running on the replicas. It goes in five steps.
- Declare. Someone with authority decides the primary site won't be back soon enough. This is a business call, written into the plan, not a technician's hunch at 2 a.m.
- Pick a recovery point. Usually the latest one. After ransomware, it's the last point before the encryption started, which is why replication history depth matters.
- Start the machines in order. Identity and DNS first, then databases, then the applications that depend on them. Orchestration tools call this a recovery plan or runbook.
- Move the users. Update DNS records, bring up VPN or remote access, and repoint public IP addresses so people and customers reach the recovered systems.
- Run in DR. You're now producing new data in the provider's cloud. That data has to come home later.
Failback is the reverse, and it's the step teams forget to plan. The provider replicates the changed data back to your repaired site, you schedule a planned failover during a quiet window, and replication resumes in the original direction. Microsoft notes that a planned failover can run "with zero-data loss," while an unplanned one loses whatever changed since the last replication.
Planned failovers are also how you test. Both Azure Site Recovery and AWS Elastic Disaster Recovery can start the replicas in an isolated network for a drill without interrupting ongoing replication.
RPO and RTO: What DRaaS Can Promise
Two numbers define any recovery setup. The recovery point objective (RPO) is how much data you can afford to lose, measured in time. The recovery time objective (RTO) is how long you can afford to be down. Our guide to RTO vs RPO goes deeper on setting both.
DRaaS is strong on RPO. Azure Site Recovery offers continuous replication for Azure and VMware VMs, and replication frequency "as low as 30 seconds for Hyper-V." That means seconds to minutes of lost data, not last night's backup.
RTO is harder, because the replication engine is only one part of it. The machines have to boot in the right order. DNS changes have to spread, which depends on the record's time-to-live. The VPN has to come up. Someone has to answer the phone and run the plan. A provider can promise that replicas will start within a set time. It can't promise that your application works at the end of it unless you've tested that exact path.
Not every system deserves the same treatment. A simple tiering keeps the DRaaS bill proportional:
| Tier | Example systems | Protection | Realistic target |
|---|---|---|---|
| 1 | Identity, line-of-business app, ERP, file shares people can't work without | DRaaS replication | Minutes of data, a few hours of downtime |
| 2 | Reporting, intranet, secondary apps | Backup with fast restore | Last backup, a day of downtime |
| 3 | Test servers, archives | Backup or rebuild from scripts | Days |
Replicate tier 1 only. The rest can wait for a restore.
Self-Service, Assisted or Managed DRaaS
DRaaS comes in three service models. The difference is who does the work on the worst day.
Self-service means you configure replication and recovery plans in a cloud tool and run the failover yourself. Azure Site Recovery and AWS Elastic Disaster Recovery are the common examples. It's the cheapest per server, and every decision on the day is yours.
Assisted means a provider helps design the setup and runs tests with you, but your team executes the failover. It suits teams with the skills and without the spare hours.
Managed means the provider owns the runbook and performs the failover to a contracted service level. It costs the most, and it's the only model where a phone call starts the recovery.
For an MSP, the choice decides what you're selling. Reselling self-service DRaaS means your technicians are the ones running the failover at night. A managed partner shifts that work but adds a margin layer. Our look at MSP backup solutions covers how providers package backup and recovery for clients.
What DRaaS Costs
DRaaS pricing has five lines, and the per-server fee is only one of them.
The first is a fee per protected machine, which licenses replication and orchestration. The second is replica storage: a full copy of every protected disk, plus the snapshots behind each recovery point you keep. The third is replication infrastructure, the small servers or appliances that receive the data.
The last two are variable. Compute for full-size recovery machines is billed only while they run, during tests and real failovers. Network costs cover egress between regions, public IP addresses, VPN gateways and bandwidth.
The two public clouds publish their list prices, which makes them a useful benchmark. As of 30 September 2026, Azure Site Recovery lists $25 per month for each VM replicated to Azure in the US East region, and every protected instance is free for its first 31 days. Microsoft states that compute for the recovery VMs "is only applied at the time of test failover and failover."
AWS Elastic Disaster Recovery lists $0.028 per hour for each source server in US East (N. Virginia), roughly $20 a month. AWS's own example for 100 servers comes to about $6,389 a month, and only $2,044 of that is the service fee. The rest is staging disks ($1,425), snapshots ($1,846.50) and replication servers ($1,053.54). In that example, the per-server fee is under a third of the bill.
Managed providers package these lines differently. In this r/msp thread, one MSP describes paying for storage plus compute time when VMs run, with some monthly compute included so drills don't start a meter. Another says their vendor charges one flat price per protected machine with everything included.
Whichever model you pick, ask how test failovers are billed. A plan that charges full compute for every drill is a plan that quietly discourages drills.
Test It or It Doesn't Count
A DRaaS setup that has never been failed over is a monthly invoice with a theory attached. Replication can show green while the recovered machines can't find a domain controller, or while a licence server refuses to start on new hardware IDs.
This r/cybersecurity post is the pattern in one story. The team had promised leadership a 4-hour RTO. Their first full test in two years took 9 hours, in a calm, controlled environment, and the runbook still referenced servers they had decommissioned.
Test failovers into an isolated network are the whole reason to buy DRaaS over plain backup, so use them. Time each step against the RTO you promised, and log what broke. Our guide to disaster recovery testing covers the five test types and when to run each.
Between tests, check that every protected server still has its replication agent running. OpenFrame can run that check as a script across a client's devices and collect the output in one place.
Questions to Ask a DRaaS Provider
Ask these before signing, and get the answers in the contract, not the sales deck.
- What RPO and RTO does the SLA commit to, and what counts as the start of the clock?
- Where do the replicas live, and in which region or country?
- How many test failovers are included each year, and how is test compute billed?
- How far back does the recovery point history go, and can it be set to a point before a ransomware infection?
- Are the replicas and snapshots immutable, and who can delete them?
- How do users and customers reach recovered systems: DNS, public IPs, VPN, remote desktop?
- During a regional disaster, is capacity reserved for us, or do we compete with every other customer?
- Who owns and updates the runbook, and who runs the failover at 2 a.m.?
- How does failback work, how long does it take, and is it included?
- If we leave, how do we get our data and replicas back?
When DRaaS Isn't Worth It
DRaaS pays off when a few servers carry the business and a day of downtime is expensive. It's the wrong purchase in four situations.
Your systems already run in SaaS. If email, files and line-of-business apps live in Microsoft 365 and cloud applications, recovering the infrastructure is the vendor's job. What you need is SaaS backup for your own data, which our upcoming guides to small business and enterprise backup cover.
A few days of downtime is acceptable. If the business can work on paper and phones for three days, a good image backup and a rebuild plan cost far less.
You have one or two servers. Many backup appliances can boot a protected server directly from the backup as a virtual machine, which covers a small office without a second site.
Nobody will test it. The money is better spent on backups you can restore than on a replica nobody has ever booted.
The Short Version
Disaster recovery as a service keeps a replicated copy of your most important servers at a provider and starts it when your site is down. Replication gives it a low RPO, while RTO still depends on boot order, networking and practice. Budget for storage, snapshots and test compute, not only the per-server fee, and run test failovers against the RTO you've promised. If you're deciding which systems earn replication and which can wait for a restore, start with the recovery tiers in your BCDR plan.
Conrad Lunderstedt
Solution Architect
I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.
