BCDR stands for business continuity and disaster recovery. It's the combined discipline of keeping a client's critical operations running during a disruption (the business continuity half) and restoring their data, systems, and infrastructure after one (the disaster recovery half). A BCDR plan is the single document that defines both: what stays up, what comes back, how fast, and who does it.
For an MSP, BCDR isn't a product you resell. It's a service line with defined recovery targets, a testing schedule, and a cost structure you either control or get squeezed by.
Most pages ranking for this term stop right before the part that decides whether your BCDR practice makes money: what it costs per endpoint, what it costs to leave a vendor, and what breaks the first time you run a real restore. That's what's below.
Business Continuity vs Disaster Recovery: The Real Difference
The two words get used interchangeably and they shouldn't be. They answer different questions and they're owned by different people.
Business continuity answers "can the client keep operating?" It covers the manual workaround when the ERP is down, the phone tree, the alternate site, the decision about which five business functions matter and which twenty can wait a week. It's a business problem. The client owns it, and most of them haven't written it down.
Disaster recovery answers "can we bring the systems back?" It covers backup, replication, failover, restore validation, and the runbook your tech follows at 2am. It's a technical problem. You own it.
The gap between the two is where MSPs get burned. You can hit a perfect four-hour RTO on the file server and still have a client dead in the water because nobody documented that their production line depends on a Windows XP box running a proprietary controller that isn't in your backup scope. Recovering IT isn't the same as recovering the business. If you sell BCDR, you're on the hook for both conversations, even when only one of them is billable.
RTO, RPO, and MTD With Real Numbers
Three metrics carry the whole discipline. Definitions are easy. Attaching numbers to them is where the work is.
RTO (recovery time objective) is how long a system can be down before the damage is unacceptable. RPO (recovery point objective) is how much data the client can afford to lose, measured in time. MTD (maximum tolerable downtime) is the outer limit for the business function itself, and it's always larger than the RTO of any single system underneath it.
Here's what those numbers look like against real workloads and what each tier costs to deliver.
| Workload | RTO | RPO | Delivery method | Rough monthly cost per client |
|---|---|---|---|---|
| Line-of-business server (accounting, ERP) | 1 hour | 15 min | Local appliance + cloud replication, instant virtualization | $300 to $500 |
| File and print server | 4 hours | 1 hour | Image backup, local restore first | $150 to $250 |
| Microsoft 365 tenant | 8 hours | 24 hours | SaaS backup, per-seat | $3 to $5 per seat |
| Domain controller | 2 hours | 4 hours | Image backup, tested bare-metal restore | $80 to $150 |
| Workstation fleet | 24 hours | 24 hours | Image backup, reimage from gold build | $2 to $6 per endpoint |
Two things fall out of that table. First, the tightest RTO tier costs roughly three times the loosest, so every hour you shave off an RTO has a price tag you can put in front of the client. Second, a 15-minute RPO and a 24-hour RPO are different products, not different settings, and clients who ask for the first while budgeting for the second need that spelled out in the proposal rather than discovered during an incident.
Vendors advertise around these numbers. Axcient's x360Recover markets a 15-minute RPO and sub-one-hour RTO, which is a fair description of what a local appliance with cloud replication does. It's not what a nightly cloud-only backup does, no matter what the datasheet says.
What Goes Into a BCDR Plan
A BCDR plan that survives contact with an auditor or an incident has nine components. Anything shorter is a backup policy wearing a costume.
- Business impact analysis. Which functions matter, ranked, with an MTD attached to each. This is a client interview, not a technical exercise.
- Asset and dependency inventory. Every system in scope, plus what it depends on. The dependency column is the one people skip and the one that ruins recoveries.
- RTO and RPO per system, signed off by the client in writing.
- Backup architecture. Where copies live, retention, immutability, air-gap status. The 3-2-1-1-0 rule is the current baseline: three copies, two media types, one offsite, one immutable, zero errors on verification.
- Recovery runbooks. Step-by-step, per system, written so a tech who didn't build it can follow it.
- Communication plan. Who calls the client, who talks to their staff, who handles insurance and legal notification. Include phone numbers that don't live in the email system you just lost.
- Testing schedule with defined test types and pass criteria.
- Roles and escalation, including what happens when your on-call tech is unreachable.
- Review cadence. Quarterly at minimum, and after every material change to the client's environment.
Get the business impact analysis right and the rest follows. Get it wrong and you'll build a technically excellent recovery for systems the client could have lived without for a week.
What BCDR Costs a 200-Endpoint MSP
Nobody on this SERP publishes real numbers, so here they are. This models a mid-sized practice: 200 endpoints across roughly 20 clients, 35 protected servers, 900 Microsoft 365 seats.
| Cost line | 2026 pricing | Annual at this scale |
|---|---|---|
| Appliance hardware | $1,500 to $2,500 (small sites), $3,000 to $8,000+ (larger) | $0 to $24,000, depending on model |
| Cloud replication per protected server | $100 to $300 per device per month | $42,000 to $126,000 |
| Workstation image backup | $2 to $6 per endpoint per month | $4,800 to $14,400 |
| Microsoft 365 backup | $3 to $5 per seat per month | $32,400 to $54,000 |
| Restore egress (variable) | $0 (Wasabi, Backblaze) to $0.09/GB (AWS) | $0 to $4,000+ |
| Tech time on testing and remediation | 6 to 10 hours per month | $7,000 to $14,000 fully loaded |
Call it $90,000 to $230,000 a year before you've billed a client. The spread is enormous, and most of it comes down to two decisions: whether your appliance hardware is capitalized or bundled, and whether your cloud tier charges egress.
Datto shifted its appliance model so SIRIS and ALTO hardware ships free-to-use across contract terms, which converts what used to be a capital purchase into monthly operating cost. Invenio IT's 2026 pricing breakdown puts a typical five-server client at $500 to $800 a month all-in through an MSP. That's a workable number if your gross margin target is 50% and you've priced the tier correctly. It's a bad number if you sold BCDR as a flat add-on three years ago and never repriced.
Egress is the line that catches people. A 500GB restore runs about $36 on AWS after the free tier, roughly $25 on Google Cloud, and $0 on Wasabi or Backblaze. On a 2TB restore that spread widens to $0 versus $175. Small in isolation. Not small when your worst client year involves nine restores and you quoted the contract assuming zero. Our breakdown of MSP backup solutions goes deeper on how the per-endpoint and per-GB models compare once restore volume is factored in.
The Exit Cost Nobody Publishes
Here's the section every vendor blog leaves out. What does it cost to move BCDR vendors?
| Exit cost component | What it involves | Typical scale |
|---|---|---|
| Data reseeding | Rebuilding full chains on the new platform from scratch | 60 to 90 days of migration work per 100 endpoints |
| Dual-running | Two agents, two consoles, two bills during cutover | 2 to 4 months of duplicate licensing |
| Appliance disposition | Buyback terms, or hardware you now own and can't repurpose | $0 to full residual value, contract dependent |
| Runbook rebuild | Every recovery procedure rewritten and retested | 1 to 2 hours per protected system |
| Contract termination | Early exit penalties on multi-year terms | Often the remaining contract value in full |
| Client communication | Explaining why the recovery report format changed | Unbillable, and it costs trust |
Industry estimates put the all-in cost of leaving an entrenched vendor at $50,000 or more for a mid-sized MSP once retraining and reduced productivity are counted. At the larger end, switching costs on MSP platform contracts reach $325,000, and BCDR is one of the stickiest lines in the stack because the data is large, the formats are proprietary, and the migration window is exactly when you're most exposed.
The practical defense is to negotiate the exit at signing, not at renewal. Three clauses matter: a documented export path in a restorable format, published egress pricing that can't change mid-term, and appliance buyback terms in writing. Vendors will give you these if you ask before you sign. Almost none will volunteer them.
This is also why more MSPs are re-examining how many of their tools sit under one contract. Platforms like OpenFrame take the AI-native all-in-one approach, with native PSA included alongside RMM and endpoint management, priced affordably and built without vendor lock-in, so the exit conversation stays a business decision rather than a hostage negotiation. It's one option among several, and the point isn't the logo on the console. It's whether your data leaves when you do. If you're mapping alternatives specifically in the backup and continuity space, our Datto alternatives guide covers the current field.
Where Open Source and Hybrid BCDR Hold Up
Open source components can carry real weight in a BCDR stack. They don't carry all of it, and pretending otherwise costs credibility with the exact technicians you're trying to convince.
Where it works: Proxmox Backup Server handles VM-level backup with deduplication and verification well enough for production, and it's genuinely free of per-device licensing. Restic and Borg are solid for file-level and server backup with encryption and immutable repository support. Bacula and UrBackup cover heterogeneous environments. Paired with cheap object storage that doesn't charge egress, the raw cost per protected server drops by an order of magnitude.
Where it breaks: multi-tenancy. Almost none of these tools were designed to isolate 20 clients under one console with per-tenant reporting, per-tenant retention policy, and role separation that survives a compliance review. You end up running separate instances per client, and the labor of maintaining 20 of anything erases the licensing savings fast. Instant virtualization, the feature that lets you boot a failed server from the backup appliance in minutes, is also thin outside commercial products. Nor do you get the vendor-supplied recovery report that half your clients' cyber insurance carriers now ask for.
The realistic pattern is hybrid: commercial BCDR on the tier-one servers where RTO is under four hours and insurance evidence matters, open source underneath it for workstation images, dev environments, and long-tail retention. That's not a compromise, it's a cost structure that reflects what each tier is worth.
Testing Cadence, and Why Most BCDR Fails Here
The uncomfortable data: roughly 22% of organizations still have no formal disaster recovery program at all, per the Disaster Recovery Journal's 2026 preparedness report. Among those that do, Secureframe's 2026 statistics roundup puts automated testing at never for 82% of setups, with only 57% of backups completing successfully and 61% of restores succeeding. One in five backups turns out to be unusable when someone finally tries it.
A backup that's never been restored is a hypothesis.
A cadence that holds up looks like this. Automated restore verification runs nightly on every protected system, checking that the image mounts and the file system is readable. A single-file or single-mailbox restore gets performed monthly per client, which takes about ten minutes and catches permission and retention drift. A full server boot test runs quarterly on every tier-one system, timed against the contracted RTO, with the result written down. An annual full failover exercise covers one client end to end, including the business continuity side: can their staff work from the alternate arrangement, do the phone numbers in the communication plan connect to a human.
Put the quarterly boot test result in the QBR deck. "Your accounting server restored in 41 minutes against a contracted 60-minute RTO" is the single most effective renewal argument in managed services, and it costs you an hour a quarter to produce.
Common Failure Modes
Six patterns account for most BCDR failures in practice.
- Backup scope drift. A new VM gets built, nobody adds it to the protection policy, and it stays invisible for eight months. Reconcile your backup inventory against your RMM asset list monthly, automatically.
- Untested bare-metal restore. Image backups that verify fine but won't boot because the driver set doesn't match replacement hardware. Only a real boot test finds this.
- Retention that doesn't match the threat. Ransomware dwell time regularly exceeds 30 days. A 14-day retention window means your clean restore point is already gone by the time you detect the intrusion.
- Credentials inside the blast radius. Backup console credentials stored in the password manager that lives on the domain you just lost. Keep an offline copy.
- Undocumented dependencies. Restoring the app server without the license server, the SQL instance, or the on-prem certificate authority. The dependency map is what turns a restore into a recovery.
- Unpriced RTO creep. The client's business grew, their tolerance for downtime dropped, and nobody repriced the tier. You're now delivering a one-hour RTO on four-hour money.
A BCDR Plan Template You Can Copy
Drop this into a document per client and fill it in. It's deliberately short, because the plans that get maintained are the ones someone can read in five minutes at 2am.
Client: [name] | Last reviewed: [date] | Owner: [tech lead]
Section 1, Business impact. Critical functions ranked 1 to 5, MTD per function, revenue impact per hour of downtime.
Section 2, Protected systems. System name, role, dependencies, RTO, RPO, backup method, retention, last successful restore test with date and elapsed time.
Section 3, Recovery sequence. Numbered order of restoration with dependency notes, so nobody brings the app server up before the database.
Section 4, Runbooks. One per tier-one system. Link out rather than inline.
Section 5, Communication. Client primary and secondary contacts with mobile numbers, your escalation chain, insurance carrier and policy number, breach counsel if applicable, and a statement of who is authorized to declare a disaster.
Section 6, Test log. Date, test type, systems covered, result, elapsed time versus contracted RTO, remediation items and their owner.
Section 7, Exit position. Current vendor, contract end date, notice period, documented export path, egress pricing, appliance buyback terms. Fill this in at signing while you still have leverage.
Section 7 is the one that separates an MSP running a BCDR practice from one renting somebody else's. Every other section describes how you recover a client. That one describes whether you can recover yourself.
Marketing Manager
Ohayo! I'm Kristina, and I'm doing good things with content, SEO, social, and community at Flamingo. Before IT, I worked as a correspondent for Ukraine's Public Broadcasting Company and have a Master's in journalism.
