A backup job that reports "Success" every night tells you one thing: something got copied somewhere. It doesn't tell you which systems are covered, how much data you'd lose, or whether anyone can restore it under pressure. This guide walks through how to build a data backup plan from scratch, one step at a time, and ends with a template you can copy into a spreadsheet today.
TL;DR
- A data backup plan is a per-system record: what gets backed up, how often, where the copies live, how long they're kept, who owns them and how restores get tested.
- Start with an inventory, then set a recovery point objective (RPO) and recovery time objective (RTO) for each system. The schedule falls out of those two numbers.
- Use 3-2-1 as the floor: three copies, two media types, one copy offsite. Add one offline or immutable copy and zero restore errors, and you get 3-2-1-1-0.
- Protect the backups from the same attacker who hits production. Separate credentials, MFA, encryption and one copy the network can't reach.
- Test restores on a calendar, name an owner for every row, and review the plan when the environment changes.
What a Data Backup Plan Covers
Three documents get mixed up here. A backup strategy is the set of rules: 3-2-1, encryption everywhere, restores tested quarterly. A backup plan applies those rules to every system you run, one row at a time. A runbook is the step-by-step restore procedure a technician follows at 2 a.m.
The plan sits in the middle, and it's the one teams skip. The strategy fits on a slide, so it gets written. The runbook gets written the first time a restore goes badly. The plan, the boring table that says "the accounting database is backed up every hour to these two places and kept for 90 days," rarely exists until an auditor or an insurer asks for it.
A backup plan also isn't a disaster recovery plan. Disaster recovery covers people, sites, failover and communications during a major outage. Backups are one input to it. If you need the wider picture, our guide on what BCDR covers walks through business continuity and disaster recovery together.
NIST's Cybersecurity Framework 2.0 puts the requirement in one line. Its outcome PR.DS-11 reads: "Backups of data are created, protected, maintained, and tested." Each of those four verbs maps to a step below. Created is steps 1 to 4. Protected is step 5. Tested is step 6. Maintained is step 7, and it's the one that quietly breaks plans a year after they're written.
Step 1: Inventory What You Would Miss
Start with a list of everything that would hurt to lose. The systems nobody thinks to list are where plans fail.
Servers and databases come to mind first. Walk further. Endpoint data that never reaches a server, like the finance lead's desktop folder. SaaS data in Microsoft 365, Google Workspace and line-of-business cloud apps. Configuration: firewall rules, switch configs, DNS zones, Group Policy, scripts in your management tools. Secrets: the password vault and the recovery keys for encrypted drives. Logs you need to keep for an investigation or a regulator.
Don't build the inventory from memory. Pull it from sources that already exist: your device inventory for servers and endpoints, the admin centers of your SaaS tenants, and the finance team's list of software subscriptions. That last one catches the apps someone bought on a credit card and never told IT about. Walk each department head through the list too, and ask one question: "What would stop you working tomorrow?"
For each item, write down where it lives, who owns the data, and what breaks if it's gone. That last column does the prioritizing for you. "Payroll can't run" and "the marketing team loses last year's banner files" are different rows.
SaaS deserves its own pass. Recycle bins, version history and retention policies in Microsoft 365 and Google Workspace are retention features with time limits. They aren't a copy you control. Our post on Microsoft 365 backup covers where Microsoft's protection stops and yours has to start.
Sync tools deserve a warning too. OneDrive, Google Drive and Dropbox copy deletions as faithfully as they copy files. That makes them great for access and useless as a backup on their own. This r/sysadmin thread from a small school district shows the pattern: a user migrates a few hundred documents, ends up with shortcuts pointing at nothing, and empties the Recycle Bin.
The replies split into two camps. One points out that recovery tools rarely help on SSDs, because freed blocks get wiped in the background. The other says backup is infrastructure you design before the loss, and software comes after that. Both answers point at the same fix: a plan written before anyone needs it.
Step 2: Set RPO and RTO for Each System
Two numbers drive every other decision in the plan.
Recovery point objective (RPO) is how much data you can afford to lose, measured in time. An RPO of one hour means that after a restore, you're missing at most the last hour of work. Recovery time objective (RTO) is how long the system can be down before the business feels serious damage. An RTO of four hours means the restore has to finish, and be verified, inside four hours.
Ask the data owner, not IT, for both numbers. IT can tell them what each number costs. A 15-minute RPO on a database means continuous or near-continuous protection, more storage and more licensing. A 24-hour RPO means one nightly job. Owners pick differently when they see the price.
Group systems into three or four tiers instead of setting unique numbers for everything. A tier model is easier to schedule, easier to test and easier to explain to the person signing the budget. Here's a common shape:
| Tier | Typical systems | RPO | RTO |
|---|---|---|---|
| Tier 1 | Line-of-business database, ERP, email | 15 min to 1 hour | 4 hours |
| Tier 2 | File servers, SaaS mailboxes and drives | 24 hours | 1 business day |
| Tier 3 | Archives, dev environments, endpoint profiles | 1 week | Best effort |
These numbers are examples, not a standard. Your owners set the real ones. Our guide to RTO vs RPO goes deeper on how to calculate them and where teams get them wrong.
Write the tier into the plan next to every system. When a new app arrives, the first question becomes "which tier?", and the schedule, retention and test cadence follow from the answer.
Step 3: Pick Your Backup Rule, 3-2-1 or 3-2-1-1-0
The 3-2-1 rule is the floor. US-CERT's paper on data backup options spells it out: keep three copies of any important file (one primary and two backups), keep the files on two different media types, and store one copy offsite.
Each number defends against a different failure. Three copies cover a single bad backup. Two media types cover a failure that hits one kind of storage, like a firmware bug across a NAS model. One offsite copy covers fire, flood, theft, or the office simply being unreachable.
3-2-1 was written before ransomware went after backups on purpose. Backup vendors have extended it since. Veeam, for example, describes 3-2-1-1-0: the same three rules, plus one copy that's offline, air-gapped or immutable, plus zero errors after recovery verification.
The extra 1 is about reach. An attacker with domain admin rights can delete anything that domain admin can reach. One copy that the network can't modify, whether it's a tape on a shelf, a rotated USB drive or object storage with a retention lock, survives that. Our post on immutable backups explains how the retention lock works and what to test before trusting it.
The 0 is about proof. A backup that has never been restored is a hope with a timestamp. Zero errors means verified restores on top of successful jobs.
This r/sysadmin thread asks the question every small team asks when they first write this down: do you have an actual policy document for 3-2-1?
One reply describes a cheap air gap: a backup NAS whose network port switches on before the job starts and off when it finishes. Others push the offsite copy to S3-compatible object storage with immutability turned on. Neither needs an enterprise budget.
Jeff Geerling's video on backup mistakes is a good gut check on your own setup before you write the rule into the plan:
Step 4: Choose Backup Types, Schedule and Retention
With tiers and a rule in place, the schedule is mostly arithmetic.
Backup type comes first. A full backup copies everything. An incremental copies what changed since the last backup of any kind. A differential copies what changed since the last full. Incrementals are small and fast but make longer restore chains. Differentials grow through the week but restore in two steps. Our post on incremental vs differential runs the storage math and the restore chains side by side.
Frequency comes from RPO. A one-hour RPO needs a job at least every hour, or continuous replication with snapshots. A 24-hour RPO needs one good job a day. If the job window can't fit inside the RPO, the RPO isn't met, whatever the dashboard says.
Retention answers a different question: how far back do you need to go? Ransomware and silent corruption can sit for weeks before anyone notices. If you keep 14 days of backups and the problem started 20 days ago, every copy you hold is already bad.
The grandfather-father-son (GFS) scheme is the usual answer. Keep daily backups for a couple of weeks, weekly backups for a couple of months, monthly backups for a year, and yearly backups for as long as your legal, contractual or insurance requirements say. That gives you fine detail for recent mistakes and coarse points for old ones, without keeping every nightly copy forever.
Set retention from outside requirements first. A regulator, a contract or a cyber insurance policy may set a minimum. Where nothing does, choose based on how long a problem could plausibly go unnoticed in your environment.
Here's how that plays out for a 25-person office with one accounting database, one file server and Microsoft 365. The database gets hourly log backups and a nightly full, because finance set a one-hour RPO. The file server gets a nightly incremental and a weekly full. Microsoft 365 gets a daily backup to a third-party service. Everything lands on a local repository first for fast restores, then copies to offsite object storage with a retention lock. Once a month, a full copy of the file server goes to an encrypted drive that's unplugged and stored offsite. That's 3-2-1-1 before a single test has run.
Write all three into the plan per system: type, frequency and retention. "Nightly incremental, weekly full, GFS 14/8/12/7" is a complete answer. "Backed up daily" isn't.
Step 5: Protect the Backups Themselves
Backups are a target. Attackers who plan to encrypt your data know you'll try to restore it, so they look for the backup server, the repository and the cloud copy first.
Sophos's State of Ransomware 2026 report puts the stakes in one number. Data was encrypted in 56% of the attacks its respondents described (2,158 organizations across 17 countries, surveyed January to March 2026). Every one of those organizations needed a clean copy the attacker didn't reach.
CISA's ransomware guide is direct about what that means: "Maintain offline, encrypted backups of critical data, and regularly test the availability and integrity of backups in a disaster recovery scenario." It also warns that "many ransomware variants attempt to find and subsequently delete or encrypt accessible backups."
Turn that into rules in the plan:
- Separate the backup credentials. The backup console and repository don't use domain accounts, and the service accounts that write backups can't delete them.
- Put MFA on every backup console and every cloud storage account that holds a copy.
- Keep the backup server off the domain, or at least out of the admin groups an attacker will reach first.
- Encrypt backups at rest and in transit, and store the encryption keys somewhere that isn't the backup server.
- Keep one copy offline or immutable, per the 3-2-1-1-0 rule above.
- Alert on backup deletions, retention changes and jobs that suddenly shrink.
The last point catches attacks in progress. A job that normally writes 400 GB and suddenly writes 4 GB, or a retention policy changed at 3 a.m., is a signal worth waking someone for. Our guide to ransomware attacks walks through how an attack moves hour by hour, including the point where it reaches backups.
Step 6: Test Restores on a Schedule
A green job status proves a copy was written. It doesn't prove the copy can be read, that the application will start from it, or that it restores inside your RTO. Only a restore proves that.
Put three kinds of test in the plan. File-level restores check that you can find and return a single item fast; run them monthly on a rotating sample. Full-system restores check that a server or VM boots and its application works; run them quarterly for every Tier 1 system. A full recovery exercise checks the people and the runbook as well as the data; run it at least once a year.
Time every test and compare it with the RTO. A restore that works but takes 14 hours against a 4-hour RTO is a failed test. Record the result in the plan: date, system, who ran it, how long it took, what broke. That record is what an auditor or insurer will ask for.
Our guide to disaster recovery testing covers the five test types, from tabletop to full failover, and how often to run each. For a shorter walkthrough of the habit itself, Ask Leo's video on testing backups covers the basics well:
Step 7: Assign Owners and Keep the Plan Current
Every row in the plan needs a named owner, and ideally two. The data owner decides RPO, RTO and retention. The technical owner runs the backups, watches the alerts and runs the tests. When those roles blur, nobody updates the plan when things change.
Things change constantly. A new SaaS app gets bought on a credit card. A file server moves to SharePoint. A database gets a second instance. Each of those should trigger a plan update, but only if someone is responsible for noticing.
Build the review into the calendar. Check the job and alert status every day, run file-restore tests monthly, run full-system tests quarterly, and review the whole plan against the inventory every year, or after any major change. A plan reviewed annually will be wrong for part of the year, and that's fine as long as the changes feed back into it.
Store the plan where you can reach it during an outage. If the only copy is on the file server you're trying to restore, you have a problem. Keep a copy in a separate system and a printed or offline copy with the runbooks.
If you run OpenFrame, you can run a script across a client's devices to report each backup agent's service state and last successful job, and see all the output in one place instead of logging into each machine.
The Data Backup Plan Template
Here's the template. Copy the columns into a spreadsheet and add one row per system. The example rows show the level of detail each cell needs; replace them with your own.
| System | Data owner / tech owner | Tier | RPO / RTO | Backup type and frequency | Copies and locations | Retention | Protection | Restore test |
|---|---|---|---|---|---|---|---|---|
| Accounting database (SQL Server) | Finance lead / IT lead | 1 | 1 hour / 4 hours | Hourly log backups, nightly full | Local repository, offsite object storage (immutable) | GFS 14 daily, 8 weekly, 12 monthly, 7 yearly | Separate backup account, MFA, encrypted | Quarterly full restore to test server, timed |
| File server | Operations manager / IT lead | 2 | 24 hours / 1 business day | Nightly incremental, weekly full | Local NAS, offsite cloud copy, monthly offline drive | GFS 14 / 8 / 12 | Backup NAS off the domain, deletion alerts | Monthly file restore, sample of 5 files |
| Microsoft 365 mail, OneDrive, SharePoint | Office manager / IT lead | 2 | 24 hours / 1 business day | Daily SaaS backup | Third-party cloud backup, separate tenant or storage | 1 year minimum | Separate admin account with MFA | Monthly mailbox item and site restore |
| Laptops (user profiles) | Each user / IT lead | 3 | 1 week / best effort | Folder redirection plus weekly endpoint backup | Cloud sync plus endpoint backup storage | 90 days | Encrypted endpoint agent | Quarterly profile restore on a spare device |
| Firewall and switch configs | IT lead / IT lead | 2 | On change / 4 hours | Export after every change, weekly scheduled export | Config repository, offsite copy | 12 months of versions | Repository access limited to IT | Annual restore to spare or lab device |
| Password vault and recovery keys | IT lead / IT lead | 1 | On change / 1 hour | Encrypted export after changes | Two offline copies in separate locations | Latest 3 versions | Encrypted, access limited to 2 people | Annual test open of each offline copy |
Before you call the plan finished, run it against this checklist:
- Every system in the inventory has a row, including SaaS, configs and secrets.
- Every row has a data owner and a technical owner by name.
- RPO and RTO come from the data owner, and the schedule meets them.
- The backup copies meet 3-2-1, and at least one copy is offline or immutable.
- Retention covers outside requirements and the time a problem could go unnoticed.
- Backup credentials are separate from domain admin, with MFA everywhere.
- Every row has a restore test with a date, a duration and a result.
- The plan itself is stored outside the systems it protects.
- Changes to the environment trigger a plan update, and someone owns noticing.
- The whole plan has a review date on the calendar.
Backup Best Practices That Keep the Plan Working
A few habits separate plans that survive contact with an incident from plans that don't.
Back up configuration as well as data. Rebuilding a server from a clean image is fast. Rebuilding its configuration from memory isn't. Firewall rules, Group Policy, DNS and your management scripts belong in the plan.
Watch for silent failures. Jobs that succeed with warnings, jobs that skip locked files, and jobs pointed at an empty or remapped volume all report green. Review the size and duration of each job against its normal range, not just its status.
Keep the restore path simple. The person restoring may not be the person who built the system. Write the runbook for the least experienced technician who might have to use it, and test it with that person.
Budget for restore time as well as storage. Pulling terabytes back from cold cloud storage can take days and cost egress fees. Check how long a full restore from each copy takes, and make sure the fastest copy covers the systems with the tightest RTO. Our guide to small business backup compares where each kind of storage fits.
Larger environments with appliances, tape and several sites face the same question at a bigger scale. Our overview of enterprise backup options covers how those pieces fit together.
Treat the plan as a living document. The first version will have gaps. The review cycle closes them.
A Plan Is a Table You Keep Current
A data backup plan isn't complicated. It's an inventory, two numbers per system, a rule for copies, a schedule, protection for the backups themselves, restore tests on a calendar, and names next to every row. The template above gives you the columns. The hard part is filling them in with the data owners and keeping them current.
If you're building this for clients, start with the Tier 1 systems and the restore tests. Those two carry most of the risk. Then read our guide on MSP backup solutions for how backup tools fit the plan you've just written.
Dmytro Koval
Head of Product Engineering
Hi! My name is Dmytro, but everyone calls me Dima. I’m a Software Developer and together with the development team, I help bring Flamingo to life — putting it on its feet from a technical perspective. Originally from Lviv, Ukraine 🇺🇦, but currently based in Spain, where I’ve been enjoying the blend of great weather, culture, and nature. I’m passionate about the mountains and love traveling — exploring new places and cultures really inspires me. These experiences constantly recharge me and give me a fresh perspective, both personally and professionally.
