Updated: October 2026
A client says "we have an SLA with you," and nobody on the team can say what it promises. Half the arguments about slow tickets start there. Here's what an SLA means in business, how it differs from an OLA and an SLO, and how to write and measure one your team can keep.
What Is an SLA?
A service level agreement (SLA) is the part of a contract that says how well a service will be delivered, not just what the service is. It turns "we'll support your IT" into numbers both sides can check: how fast someone responds, how fast a fix lands, how much uptime counts as normal.
A usable SLA answers seven questions:
- Scope. Which services, users and devices are covered, and which aren't.
- Service hours. When the clock runs: business hours, extended hours or 24/7.
- Priorities. How a ticket gets its P1 to P4 label.
- Targets. Response and resolution times per priority, plus any uptime target.
- Measurement. Which system records the times, and what pauses the clock.
- Remedies. What happens when a target is missed: service credits, a review, an exit clause.
- Review. How often both sides look at the numbers and change them.
If one of those is missing, the gap gets filled during an outage, by whoever is angriest.
SLA vs OLA vs Underpinning Contract
An SLA is the promise to the customer. Keeping it usually depends on people who never signed it.
An operational level agreement (OLA) is an internal agreement between teams. It "defines interdependent relationships in support of a service-level agreement," in Wikipedia's wording of the ITIL idea. If the service desk promises a four-hour fix for a P2, the network team needs an OLA that says it will pick up an escalation within one hour.
An underpinning contract (UC) does the same job with an outside supplier: the ISP, the hardware vendor, the SaaS platform. When the internet line is down, your resolution time is only as good as the carrier's repair commitment.
Write the three in order. Set the customer SLA, then check every internal team and supplier it depends on can hit its share. An SLA with no OLA or UC behind it is a guess.
SLA vs SLO vs SLI
These three sound alike and get used interchangeably. Google's Site Reliability Engineering book pins them down.
An SLI is "a carefully defined quantitative measure of some aspect of the level of service," like the share of successful logins or median ticket response time. An SLO is "a target value or range of values for a service level that is measured by an SLI," like 95% of P2 tickets answered within an hour. An SLA is "an explicit or implicit contract with your users that includes consequences of meeting (or missing) the SLOs they contain."
So the SLI is the measurement, the SLO is the goal, and the SLA is the promise with consequences attached. Teams often run internal SLOs a notch tighter than the SLA, so they get warned before the contract is breached.
Google Cloud's short explainer walks through the same split with examples:
Response Time vs Resolution Time
Response time is how long until a person acknowledges the ticket and starts work. Resolution time is how long until the problem is fixed and the user confirms it. They measure different things, and SLAs that blur them cause fights.
Response is fully in your control, so it's easy to promise. Resolution isn't. You can't guarantee a fix when you don't yet know the cause, or when the fix depends on the carrier. That's why a common pattern is to commit hard to response and treat resolution as a target with named exceptions.
This r/sysadmin thread on P1 and P2 targets lands on the same split. The top reply calls resolution "a hard thing to put on an SLA" and suggests a one-hour maximum to start work on a P1:
Two rules decide whether the numbers mean anything. First, say which clock runs: a four-hour target on business hours can stretch over a weekend, while a 24/7 clock can't. Second, list what stops the clock. "Waiting on customer" and "waiting on third party" usually pause it, and the SLA should say so in writing.
A Sample SLA Table
Priority comes from impact and urgency: how many people are affected, and how fast the damage grows. Our guide to the incident management process covers the priority matrix in full. Once priorities exist, the targets fit in one table:
| Priority | Example | Response | Resolution target | Clock |
|---|---|---|---|---|
| P1 Critical | Whole office can't work, server or line down | 30 minutes | 4 hours | 24/7 |
| P2 High | A team or a key app is down, workaround exists | 1 hour | 8 business hours | Business hours |
| P3 Normal | One user affected, can keep working | 4 business hours | 3 business days | Business hours |
| P4 Low | Requests, how-to questions, new setups | 1 business day | 5 business days | Business hours |
These numbers are a starting point, not a standard. The right target is what the client's downtime costs, weighed against what faster cover costs you. Tighter numbers need more people on call, and that belongs in the price.
Service Credits and Penalties
The consequence is what makes an SLA more than a wish. The most common form is a service credit: a discount on the next invoice when a target is missed.
Cloud providers publish theirs. The Amazon EC2 SLA (last updated May 2022) commits to 99.99% monthly uptime for instances spread across availability zones. Drop below that and the credit is 10% of the bill, rising to 30% below 99.0% and 100% below 95.0%. That 99.99% leaves room for about 4.3 minutes of downtime in a 30-day month.
For a service desk, credits usually attach to response times, because those are fully in your hands. Keep the remedy proportional. A credit that wipes out a month's margin for one missed P3 makes the team defensive, and defensive teams game metrics.
Measuring SLAs Without Gaming Them
Measure SLAs in the PSA or ticketing system, not in a spreadsheet someone fills in on Friday. The system stamps created, first response, status changes and resolved times, and the report comes straight from those stamps.
Then watch for the ways the numbers get bent. An auto-reply that counts as "first response." Tickets parked in "waiting on customer" with no question asked. P2s quietly downgraded to P3. Tickets closed and reopened to restart the clock. Each one keeps the dashboard green while users wait.
Pair the SLA numbers with signals that are harder to fake: reopen rate, customer satisfaction on closed tickets, and a monthly sample of breached tickets read by a human. This r/msp debate on "blind adherence" is a good reminder that an SLA is a floor. The top answer: when people are free, low-priority work gets done right away.
SLAs in Short
An SLA turns a service into numbers: response and resolution times per priority, the clock they run on, what pauses it, and what happens when a target is missed. OLAs and supplier contracts are what make those numbers keepable, and SLOs give the team an early warning. Measure it in the ticketing system and pair it with signals nobody can game.
If you're on the buying side, our list of questions to ask an MSP covers how to read an SLA before you sign.
Conrad Lunderstedt
Solution Architect
I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.
