Downtime is easy to complain about and hard to put a price on. Without that number, backup budgets, monitoring tools and maintenance windows all get argued on gut feel. This guide explains downtime meaning in IT terms, how to measure it, and how to put a price on it with a formula you can fill in yourself.
TL;DR
- Downtime is any period when a system can't do its job for the people who depend on it. A server that's switched on but too slow to use counts too.
- There are two kinds. Planned downtime is maintenance you schedule and announce. Unplanned downtime is everything else, and it's the kind that costs money.
- Measure it as availability. Uptime percentage = (total time minus downtime) divided by total time. 99.9% sounds high and still allows about 43 minutes of downtime a month.
- Price it with a formula, not a headline stat. Lost productivity + lost revenue + recovery cost + penalties. A 30-person firm with a four-hour outage lands around $8,400 in the worked example below.
- SLA credits won't cover it. Provider credits are a slice of the monthly fee, not a refund of what the outage cost you.
What Does Downtime Mean?
In everyday English, downtime means a break. Time off, a quiet afternoon, a pause between jobs. That's why dictionaries own the top of the search results for this word.
In IT, downtime means a period when a system is unavailable to the people who need it. The email server doesn't accept mail. The accounting app won't load. The office internet drops and every cloud tool goes with it.
The key phrase is "to the people who need it". A server can be powered on, answering pings and showing green on a dashboard while users stare at a spinning wheel. If they can't work, that's downtime.
That's why IT teams split availability into three states:
- Up: the system does its job at normal speed.
- Degraded: it works, but slowly or partly. Some users can log in, some can't. Mail arrives an hour late.
- Down: it doesn't work at all for the people who depend on it.
Degraded time is the tricky one. Monitoring tools often count it as uptime because the service still responds. Users count it as downtime because they can't get their work done. When you report downtime, say which definition you're using.
An outage and downtime are close cousins. "Outage" usually describes the event, like a power cut or a failed update. "Downtime" describes the time lost because of it. One outage can cause downtime for several systems at once.
Planned vs Unplanned Downtime
Not all downtime is a failure. Some of it is on the calendar.
Planned downtime is maintenance you schedule on purpose. Patching servers, replacing a switch, upgrading a line-of-business app, migrating mail. You pick the time, you warn users, and you have a rollback plan. A Saturday-night window where the file server reboots for updates is planned downtime.
Unplanned downtime is everything you didn't choose. A disk fails. An update breaks the VPN client. Ransomware encrypts the file share. A cloud provider has a bad day. You find out when the phones start ringing.
The difference matters for cost. Planned downtime happens when few people are working, and they know it's coming. Unplanned downtime hits in the middle of a Tuesday, with the whole office waiting and nobody sure how long it will last.
It also matters for contracts. Service level agreements usually exclude planned maintenance from the downtime they count. Microsoft's guide to reading SLAs lists planned maintenance among the common exclusions, alongside customer misconfiguration and features still in preview. If your provider schedules a window, their SLA clock often stops, even though your users still can't work.
The goal for IT is simple: move as much downtime as possible from the unplanned column to the planned one. A patch installed during a scheduled window is ten minutes of planned downtime. The same patch skipped for six months can become a day of unplanned downtime.
Degraded service sits in a grey zone of its own. An r/msp thread titled simply "Service degradation on Microsoft 365" is one of many like it.
What Causes Unplanned Downtime
The causes fall into a few buckets, and people sit in the middle of several of them.
Human error and process. Someone skips a step, changes the wrong firewall rule, or restarts the wrong server. Uptime Institute's 2025 outage analysis found that nearly 40% of organizations had a major outage caused by human error over the previous three years. Most of those came from staff not following procedures, or from procedures that were flawed to begin with.
Hardware. Disks, power supplies, switches and UPS batteries wear out. They rarely fail at a convenient time.
Software and updates. A bad patch, a driver conflict, an expired certificate, a database that fills its disk.
Security incidents. Ransomware, account takeover, DDoS attacks. These tend to produce the longest downtime because recovery has to wait for the investigation.
Third parties. Your ISP, your cloud provider, your SaaS vendors, your DNS host. When they go down, you go down, and there's little you can fix yourself.
That last bucket catches teams out. In November 2025, a Cloudflare outage took down sites that had nothing wrong with their own servers. One sysadmin described every alert turning red while CPU, logs and health checks all looked fine. The team spent the morning restarting services and rolling back deploys before they found the cause upstream.
The lesson from that thread is about perspective. Check your systems from outside as well as inside, so you can tell "our server is down" apart from "the path to our server is down".
How to Measure Downtime
You can't reduce what you don't count. Three numbers cover it.
Availability, or uptime percentage. The share of time a system was usable over a period:
Availability = (Total time - Downtime) / Total time x 100
Over a 30-day month there are 43,200 minutes. If the file server was down for 90 minutes, availability was (43,200 - 90) / 43,200, or 99.79%.
Mean time to detect (MTTD). How long it takes from the moment something breaks to the moment someone knows. If users report outages before your monitoring does, this number is the first one to fix.
Mean time to recover (MTTR). How long from detection to the service working again. Recovery includes diagnosis, the fix, and checking that users can work. Our guide to the incident management process covers how priorities and escalation shape this number.
Uptime percentages hide a lot in their decimals. Each extra nine cuts the allowed downtime by a factor of ten:
| Availability | Downtime per 30-day month | Downtime per year |
|---|---|---|
| 99% | 7 hours 12 minutes | 3 days 15 hours 36 minutes |
| 99.5% | 3 hours 36 minutes | 1 day 19 hours 48 minutes |
| 99.9% | 43 minutes 12 seconds | 8 hours 45 minutes 36 seconds |
| 99.95% | 21 minutes 36 seconds | 4 hours 22 minutes 48 seconds |
| 99.99% | 4 minutes 19 seconds | 52 minutes 34 seconds |
| 99.999% | 26 seconds | 5 minutes 15 seconds |
Two cautions before you promise a number. First, the percentage only means something with a definition of downtime next to it. Microsoft's guide to reading SLAs points out that a 99.9% SLA doesn't mean you should expect only 8.7 hours of downtime a year, because the provider's definition decides what gets counted. Second, a business service depends on a chain of systems. The accounting app needs the server, the network, the internet line and the identity provider. Each link has its own availability, and users feel the weakest one.
If you're writing these targets into a contract, read our explainer on what an SLA is first.
How Much Does Downtime Cost?
The headline numbers are big. ITIC's 2024 survey of more than 1,000 firms found that one hour of downtime costs over $300,000 for more than 90% of mid-size and large enterprises.
A June 2024 study by Splunk and Oxford Economics put the cost of downtime for Global 2000 companies at 9% of profits.
Both studies describe large enterprises. A 30-person accounting firm doesn't lose $300,000 an hour. Quote an enterprise stat to a small business owner and they'll stop listening.
A better approach is to price your own downtime. The math takes ten minutes and gives you a number your client or your finance team will recognize.
UptimeRobot's July 2026 video walks through where these cost numbers come from, which costs people forget, and a worked example of its own:
The Downtime Cost Formula
Downtime cost has four parts. Add them up for one outage, then divide by the hours to get a cost per hour.
Cost = Lost productivity + Lost revenue + Recovery cost + Penalties
Lost productivity. People affected x their loaded hourly cost x the share of their work that stopped x hours down. Loaded cost means salary plus benefits and overhead, not the salary alone. Be fair about the share: some people can keep working offline, others can't do anything.
Lost revenue. Revenue per hour x the share that's lost for good x hours down. A law firm can often bill the same work tomorrow, so little revenue is lost. An online shop that's down during a sale loses orders that won't come back.
Recovery cost. Technician hours, emergency vendor fees, overtime for staff catching up, replacement hardware, data restore costs.
Penalties. SLA credits you owe your own customers, late fees, contract penalties, regulatory fines if data was involved.
There's a fifth category you can't put a number on: reputation, a client who quietly starts looking at other providers, staff morale after the third outage this quarter. Name it in the report. Don't invent a figure for it.
Worked Example: A 30-Person Firm, Four Hours Down
Here's the formula with sample inputs. These are illustrative assumptions, not benchmarks. Swap in your own.
A 30-person professional services firm loses its file server and line-of-business app for four hours on a weekday morning.
- Lost productivity: 30 people x $55 loaded hourly cost x 70% of work blocked x 4 hours = $4,620
- Lost revenue: the firm bills about $600 per working hour, and half of what those four hours would have produced can't be recovered later. $600 x 4 x 50% = $1,200
- Recovery cost: 6 technician hours at $150 = $900, plus 10 staff working 2 hours of overtime at time and a half ($82.50) to catch up = $1,650. Total $2,550
- Penalties: none in this case, so $0
Total: $8,370 for the outage, or about $2,090 per hour.
Now you have a number to plan with. If a second server with failover costs $4,000 a year and prevents one outage like this, it pays for itself twice over. If a monitoring tool cuts detection time from 45 minutes to 5, you can price those 40 minutes too.
Run the same math for each critical system. The line-of-business app, email, the internet connection and the phone system usually cost very different amounts per hour. That list tells you where to spend first. It also turns a vague "we need better IT" conversation into a budget line, the same way pricing out IT support for a small business does.
Why SLA Credits Won't Cover the Bill
When a provider misses its SLA, you usually get a service credit. That feels like compensation. It rarely comes close.
Credits are a percentage of the monthly fee for the affected service. Microsoft's guide says it plainly: if your monthly spend on a service is modest, a 10% credit might amount to a small sum, and credits don't cover lost revenue, customer attrition or reputational damage.
Say the firm above runs its line-of-business app as SaaS at $12 per seat for 30 seats. That's $360 a month. A 10% credit is $36. The outage cost $8,370.
Credits also aren't automatic. Providers generally expect you to detect the breach, document it and file a claim within a deadline, often 30 to 60 days after the billing month ends. Without your own monitoring data, you can't prove anything.
Read an SLA as a signal of how a provider expects to fail. The resilience has to come from your side: backups, redundancy, monitoring and a plan.
How to Reduce Downtime
Every outage has the same timeline. Something breaks, someone notices, someone diagnoses it, someone fixes it, and users get back to work. You can shrink each stage.
Shrink detection time with monitoring. Alerts should reach IT before users pick up the phone. Watch the things users feel, like login, the app's front page and the internet line, not only CPU and disk. Check from outside your network too, so a provider outage doesn't look like your own failure. Our guide to IT infrastructure monitoring covers what to watch.
Shrink diagnosis time with runbooks. A one-page note per critical system: what it depends on, where the logs live, who the vendor contact is, and the last three things that broke it. At 8 a.m. with the CEO waiting, nobody wants to rediscover this.
Shrink repair time with backups you've tested. Restore each backup at least once before you count on it. Set a recovery time objective for each system and test against it. Our comparison of RTO vs RPO explains how to set both.
Remove whole outages with redundancy. A second internet line from a different provider. Two domain controllers instead of one. A spare switch on the shelf. Redundancy costs money every month, which is exactly why the cost-per-hour number matters.
Move work into planned windows. Patch on a schedule, announce maintenance, and keep a rollback plan. Planned downtime at 10 p.m. on a Saturday beats unplanned downtime at 10 a.m. on a Monday.
Rehearse the big one. For security incidents, recovery waits on investigation. A written and practiced incident response plan keeps that wait short.
If your fleet runs on OpenFrame, you can run a script across a client's devices, such as a last-boot-time or service status check, and collect every device's output in one place.
Downtime, in Short
Downtime is time when a system can't do its job for the people who need it. Split it into planned and unplanned, measure it as availability, and watch detection and recovery times. Then price it with the four-part formula, using your own headcount and revenue instead of an enterprise headline.
That cost-per-hour number turns every backup, monitoring and redundancy decision into simple arithmetic. For the planning side of outages that last longer than an afternoon, read our guide to BCDR for MSPs.
Content Marketing Lead
Ohayo! I run content, SEO, social, and community at Flamingo. Before IT, I worked as a correspondent for Ukraine's Public Broadcasting Company and have a Master's in journalism.
