Flamingo Raises $4.5M Seed Round

Skip to content

Downtime is easy to complain about and hard to put a price on. Without that number, backup budgets, monitoring tools and maintenance windows all get argued on gut feel. This guide explains downtime meaning in IT terms, how to measure it, and how to put a price on it with a formula you can fill in yourself.

TL;DR

  • Downtime is any period when a system can't do its job for the people who depend on it. A server that's switched on but too slow to use counts too.
  • There are two kinds. Planned downtime is maintenance you schedule and announce. Unplanned downtime is everything else, and it's the kind that costs money.
  • Measure it as availability. Uptime percentage = (total time minus downtime) divided by total time. 99.9% sounds high and still allows about 43 minutes of downtime a month.
  • Price it with a formula, not a headline stat. Lost productivity + lost revenue + recovery cost + penalties. A 30-person firm with a four-hour outage lands around $8,400 in the worked example below.
  • SLA credits won't cover it. Provider credits are a slice of the monthly fee, not a refund of what the outage cost you.

What Does Downtime Mean?

In everyday English, downtime means a break. Time off, a quiet afternoon, a pause between jobs. That's why dictionaries own the top of the search results for this word.

In IT, downtime means a period when a system is unavailable to the people who need it. The email server doesn't accept mail. The accounting app won't load. The office internet drops and every cloud tool goes with it.

The key phrase is "to the people who need it". A server can be powered on, answering pings and showing green on a dashboard while users stare at a spinning wheel. If they can't work, that's downtime.

That's why IT teams split availability into three states:

  • Up: the system does its job at normal speed.
  • Degraded: it works, but slowly or partly. Some users can log in, some can't. Mail arrives an hour late.
  • Down: it doesn't work at all for the people who depend on it.

Degraded time is the tricky one. Monitoring tools often count it as uptime because the service still responds. Users count it as downtime because they can't get their work done. When you report downtime, say which definition you're using.

An outage and downtime are close cousins. "Outage" usually describes the event, like a power cut or a failed update. "Downtime" describes the time lost because of it. One outage can cause downtime for several systems at once.

Planned vs Unplanned Downtime

Not all downtime is a failure. Some of it is on the calendar.

Planned downtime is maintenance you schedule on purpose. Patching servers, replacing a switch, upgrading a line-of-business app, migrating mail. You pick the time, you warn users, and you have a rollback plan. A Saturday-night window where the file server reboots for updates is planned downtime.

Unplanned downtime is everything you didn't choose. A disk fails. An update breaks the VPN client. Ransomware encrypts the file share. A cloud provider has a bad day. You find out when the phones start ringing.

The difference matters for cost. Planned downtime happens when few people are working, and they know it's coming. Unplanned downtime hits in the middle of a Tuesday, with the whole office waiting and nobody sure how long it will last.

It also matters for contracts. Service level agreements usually exclude planned maintenance from the downtime they count. Microsoft's guide to reading SLAs lists planned maintenance among the common exclusions, alongside customer misconfiguration and features still in preview. If your provider schedules a window, their SLA clock often stops, even though your users still can't work.

The goal for IT is simple: move as much downtime as possible from the unplanned column to the planned one. A patch installed during a scheduled window is ten minutes of planned downtime. The same patch skipped for six months can become a day of unplanned downtime.

Degraded service sits in a grey zone of its own. An r/msp thread titled simply "Service degradation on Microsoft 365" is one of many like it.

What Causes Unplanned Downtime

The causes fall into a few buckets, and people sit in the middle of several of them.

Human error and process. Someone skips a step, changes the wrong firewall rule, or restarts the wrong server. Uptime Institute's 2025 outage analysis found that nearly 40% of organizations had a major outage caused by human error over the previous three years. Most of those came from staff not following procedures, or from procedures that were flawed to begin with.

Hardware. Disks, power supplies, switches and UPS batteries wear out. They rarely fail at a convenient time.

Software and updates. A bad patch, a driver conflict, an expired certificate, a database that fills its disk.

Security incidents. Ransomware, account takeover, DDoS attacks. These tend to produce the longest downtime because recovery has to wait for the investigation.

Third parties. Your ISP, your cloud provider, your SaaS vendors, your DNS host. When they go down, you go down, and there's little you can fix yourself.

That last bucket catches teams out. In November 2025, a Cloudflare outage took down sites that had nothing wrong with their own servers. One sysadmin described every alert turning red while CPU, logs and health checks all looked fine. The team spent the morning restarting services and rolling back deploys before they found the cause upstream.

The lesson from that thread is about perspective. Check your systems from outside as well as inside, so you can tell "our server is down" apart from "the path to our server is down".

How to Measure Downtime

You can't reduce what you don't count. Three numbers cover it.

Availability, or uptime percentage. The share of time a system was usable over a period:

Availability = (Total time - Downtime) / Total time x 100

Over a 30-day month there are 43,200 minutes. If the file server was down for 90 minutes, availability was (43,200 - 90) / 43,200, or 99.79%.

Mean time to detect (MTTD). How long it takes from the moment something breaks to the moment someone knows. If users report outages before your monitoring does, this number is the first one to fix.

Mean time to recover (MTTR). How long from detection to the service working again. Recovery includes diagnosis, the fix, and checking that users can work. Our guide to the incident management process covers how priorities and escalation shape this number.

Uptime percentages hide a lot in their decimals. Each extra nine cuts the allowed downtime by a factor of ten:

AvailabilityDowntime per 30-day monthDowntime per year
99%7 hours 12 minutes3 days 15 hours 36 minutes
99.5%3 hours 36 minutes1 day 19 hours 48 minutes
99.9%43 minutes 12 seconds8 hours 45 minutes 36 seconds
99.95%21 minutes 36 seconds4 hours 22 minutes 48 seconds
99.99%4 minutes 19 seconds52 minutes 34 seconds
99.999%26 seconds5 minutes 15 seconds

Two cautions before you promise a number. First, the percentage only means something with a definition of downtime next to it. Microsoft's guide to reading SLAs points out that a 99.9% SLA doesn't mean you should expect only 8.7 hours of downtime a year, because the provider's definition decides what gets counted. Second, a business service depends on a chain of systems. The accounting app needs the server, the network, the internet line and the identity provider. Each link has its own availability, and users feel the weakest one.

If you're writing these targets into a contract, read our explainer on what an SLA is first.

How Much Does Downtime Cost?

The headline numbers are big. ITIC's 2024 survey of more than 1,000 firms found that one hour of downtime costs over $300,000 for more than 90% of mid-size and large enterprises.

A June 2024 study by Splunk and Oxford Economics put the cost of downtime for Global 2000 companies at 9% of profits.

Both studies describe large enterprises. A 30-person accounting firm doesn't lose $300,000 an hour. Quote an enterprise stat to a small business owner and they'll stop listening.

A better approach is to price your own downtime. The math takes ten minutes and gives you a number your client or your finance team will recognize.

UptimeRobot's July 2026 video walks through where these cost numbers come from, which costs people forget, and a worked example of its own:

The Downtime Cost Formula

Downtime cost has four parts. Add them up for one outage, then divide by the hours to get a cost per hour.

Cost = Lost productivity + Lost revenue + Recovery cost + Penalties

Lost productivity. People affected x their loaded hourly cost x the share of their work that stopped x hours down. Loaded cost means salary plus benefits and overhead, not the salary alone. Be fair about the share: some people can keep working offline, others can't do anything.

Lost revenue. Revenue per hour x the share that's lost for good x hours down. A law firm can often bill the same work tomorrow, so little revenue is lost. An online shop that's down during a sale loses orders that won't come back.

Recovery cost. Technician hours, emergency vendor fees, overtime for staff catching up, replacement hardware, data restore costs.

Penalties. SLA credits you owe your own customers, late fees, contract penalties, regulatory fines if data was involved.

There's a fifth category you can't put a number on: reputation, a client who quietly starts looking at other providers, staff morale after the third outage this quarter. Name it in the report. Don't invent a figure for it.

Worked Example: A 30-Person Firm, Four Hours Down

Here's the formula with sample inputs. These are illustrative assumptions, not benchmarks. Swap in your own.

A 30-person professional services firm loses its file server and line-of-business app for four hours on a weekday morning.

  • Lost productivity: 30 people x $55 loaded hourly cost x 70% of work blocked x 4 hours = $4,620
  • Lost revenue: the firm bills about $600 per working hour, and half of what those four hours would have produced can't be recovered later. $600 x 4 x 50% = $1,200
  • Recovery cost: 6 technician hours at $150 = $900, plus 10 staff working 2 hours of overtime at time and a half ($82.50) to catch up = $1,650. Total $2,550
  • Penalties: none in this case, so $0

Total: $8,370 for the outage, or about $2,090 per hour.

Now you have a number to plan with. If a second server with failover costs $4,000 a year and prevents one outage like this, it pays for itself twice over. If a monitoring tool cuts detection time from 45 minutes to 5, you can price those 40 minutes too.

Run the same math for each critical system. The line-of-business app, email, the internet connection and the phone system usually cost very different amounts per hour. That list tells you where to spend first. It also turns a vague "we need better IT" conversation into a budget line, the same way pricing out IT support for a small business does.

Why SLA Credits Won't Cover the Bill

When a provider misses its SLA, you usually get a service credit. That feels like compensation. It rarely comes close.

Credits are a percentage of the monthly fee for the affected service. Microsoft's guide says it plainly: if your monthly spend on a service is modest, a 10% credit might amount to a small sum, and credits don't cover lost revenue, customer attrition or reputational damage.

Say the firm above runs its line-of-business app as SaaS at $12 per seat for 30 seats. That's $360 a month. A 10% credit is $36. The outage cost $8,370.

Credits also aren't automatic. Providers generally expect you to detect the breach, document it and file a claim within a deadline, often 30 to 60 days after the billing month ends. Without your own monitoring data, you can't prove anything.

Read an SLA as a signal of how a provider expects to fail. The resilience has to come from your side: backups, redundancy, monitoring and a plan.

How to Reduce Downtime

Every outage has the same timeline. Something breaks, someone notices, someone diagnoses it, someone fixes it, and users get back to work. You can shrink each stage.

Shrink detection time with monitoring. Alerts should reach IT before users pick up the phone. Watch the things users feel, like login, the app's front page and the internet line, not only CPU and disk. Check from outside your network too, so a provider outage doesn't look like your own failure. Our guide to IT infrastructure monitoring covers what to watch.

Shrink diagnosis time with runbooks. A one-page note per critical system: what it depends on, where the logs live, who the vendor contact is, and the last three things that broke it. At 8 a.m. with the CEO waiting, nobody wants to rediscover this.

Shrink repair time with backups you've tested. Restore each backup at least once before you count on it. Set a recovery time objective for each system and test against it. Our comparison of RTO vs RPO explains how to set both.

Remove whole outages with redundancy. A second internet line from a different provider. Two domain controllers instead of one. A spare switch on the shelf. Redundancy costs money every month, which is exactly why the cost-per-hour number matters.

Move work into planned windows. Patch on a schedule, announce maintenance, and keep a rollback plan. Planned downtime at 10 p.m. on a Saturday beats unplanned downtime at 10 a.m. on a Monday.

Rehearse the big one. For security incidents, recovery waits on investigation. A written and practiced incident response plan keeps that wait short.

If your fleet runs on OpenFrame, you can run a script across a client's devices, such as a last-boot-time or service status check, and collect every device's output in one place.

Downtime, in Short

Downtime is time when a system can't do its job for the people who need it. Split it into planned and unplanned, measure it as availability, and watch detection and recovery times. Then price it with the four-part formula, using your own headcount and revenue instead of an enterprise headline.

That cost-per-hour number turns every backup, monitoring and redundancy decision into simple arithmetic. For the planning side of outages that last longer than an afternoon, read our guide to BCDR for MSPs.

Kristina Shkriabina

Content Marketing Lead

Ohayo! I run content, SEO, social, and community at Flamingo. Before IT, I worked as a correspondent for Ukraine's Public Broadcasting Company and have a Master's in journalism.

Related Content

Blog Posts

Product Releases

Podcasts

Webinars

Case Studies

Events

Onboarding Guides

Frequently Asked Questions

MSP AI Agents

On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.
Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.
In the cloud, on US soil. Your data stays stateside.
An outage is the event, such as a power cut, a failed update or a provider incident. Downtime is the time a system can't be used because of it. One outage can cause downtime for several systems at once, and the downtime often lasts longer than the outage because users and data need time to recover.
It depends on what an hour of downtime costs you. 99.9% availability allows about 43 minutes of downtime in a 30-day month, and 99.99% allows about 4 minutes. Price each critical system first, then set a target per system instead of one number for everything.
Add four parts for one outage: lost productivity (people affected x loaded hourly cost x share of work blocked x hours), lost revenue that can't be earned back later, recovery cost (technician time, overtime, vendor fees, hardware) and penalties. Divide the total by the hours down to get a cost per hour.
For your users, yes: they can't work while the system is down. In provider SLAs, planned maintenance is often excluded from the downtime that counts against the uptime commitment. Check the SLA's definition of downtime before you compare uptime numbers.
Mean time to detect (MTTD) measures how long it takes from the moment something breaks to the moment someone knows. Mean time to recover (MTTR) measures how long it takes from detection until users can work again. Monitoring shortens MTTD; runbooks, tested backups and redundancy shorten MTTR.