Every RMM and PSA ships a dashboard, and every dashboard has forty tiles. Nobody reads forty tiles, so the numbers stop changing decisions and start decorating a login page. This guide picks the IT metrics worth tracking for an internal team and for an MSP reporting to clients, and shows where each one is measured, and where it gets gamed.
What IT Metrics Are (and Which Ones Are KPIs)
A metric is any number you can count from a system: tickets opened, minutes to first reply, endpoints missing a patch. A KPI is a metric with three things bolted on: an owner, a target, and a decision that changes when the number crosses the target. Ticket volume is a metric. Tickets per user per month for one client, trending up for a quarter, with a review booked when it passes 1.5, is a KPI.
Two more distinctions save arguments later. Leading metrics predict pain: patch lag, backlog growth, reopen rate. Lagging metrics record it: outages, churn, breach count. And an SLA is a promise built on a metric, so measure the metric for a quarter before you sign the promise. The post on what an SLA is covers how the promise gets written.
One rule from NinjaOne's IT KPI guide (updated April 2026) holds up in practice: five to ten KPIs per function. Past that, people watch the dashboard instead of the fleet. The rest of this post sorts the candidates into four layers: service desk, uptime, security, and what the business sees.
Service Desk Metrics
The desk produces the most numbers and the most arguments about them. Start with the ones that describe what a user feels:
- First response time: minutes from ticket creation to a human reply, measured per priority, never as one blended average.
- Resolution time: creation to confirmed fix, again per priority.
- First-tier resolution rate: the share of tickets closed by the first tier without escalation. One r/msp owner prefers it over first-contact resolution for a plain reason: a second call is fine, an escalation costs a senior tech's hour.
- Reopen rate: tickets reopened within seven days of closure. It is the check on resolution time, because a fast close that comes back was not a fix.
- Backlog trend: open minus closed over a rolling seven days. Negative is good, positive for three days in a row means the queue is winning.
- CSAT: a one-tap rating at closure, read next to its response rate, since a 4.9 from six people says less than a 4.4 from sixty.
- Tickets per user per month: the demand number, and the first thing to compare across clients or departments.
MSPs add one more: reactive hours per endpoint per month, which ties desk time to what each client pays for. A client at 0.4 hours per endpoint and a client at 1.6 are not the same account, whatever the contract says. The help desk KPIs section of our help desk guide covers the formulas in more detail.
Uptime and Reliability Metrics
Availability is a percentage per service per month, measured where the users are. A synthetic check that logs into the app every five minutes counts; the server's own uptime counter does not, because a server can be up while the service behind it is dead. The arithmetic matters when you write the target: 99.9% allows about 43 minutes of downtime a month, 99.5% allows about 3.6 hours.
Behind availability sit the repair metrics. Mean time to resolve, by priority, tells you how long an outage lasts once someone is working it. Mean time between failures on a device class tells you when hardware is aging out. Both need an incident record with real timestamps, which is what the incident management process exists to produce.
Two more belong on every scorecard because they predict the outage rather than report it. Patch compliance is the share of endpoints on the current patch level inside the agreed window, with the exceptions named. Backup health is two numbers: backup job success rate, and the date of the last successful restore test. The second one is a date, not a percentage, on purpose. Uptime Institute's 2026 outage analysis (May 2026) found 57% of respondents saying their most recent major outage cost more than $100,000, and one in five over $1 million. A restore test date is the cheapest insurance on that list. Our guide to IT infrastructure monitoring covers where the availability and repair numbers come from.
Security Metrics
Security metrics come in two kinds: coverage and time. Coverage numbers answer "how much of the estate is protected": MFA coverage on accounts, endpoint agent coverage, disk encryption coverage. Each is a percentage of in-scope devices or accounts, and each is only useful with the exceptions listed by name. "97% MFA" hides three service accounts with domain admin rights.
Time numbers answer "how fast do we close the gap". Time to patch a known-exploited vulnerability is the one to start with. CISA's Binding Operational Directive 22-01 gives federal agencies two weeks to remediate entries in the Known Exploited Vulnerabilities catalog (six months for CVEs dated before 2021); borrow that clock even if the directive does not apply to you. Verizon's 2025 Data Breach Investigations Report (April 2025) put exploitation of vulnerabilities up 34% as an initial access vector, which is the reason that clock matters. Add open vulnerabilities older than 30 days, and mean time to detect, the dwell-time number that threat hunting is meant to shrink.
Then the people numbers: phishing simulation click rate and, more useful, report rate. A team that reports the phish inside five minutes has a working control; a team that never clicks may just have a good spam filter. The post on security awareness training explains why the report rate is the one to track.
Last, the noise number. Alerts per technician per day, and the share that led to an action, tells you whether the security stack is producing signal. Tuning starts from that ratio, which the alert fatigue post walks through.
What Clients Want to Hear
A January 2026 r/msp thread asked which KPIs clients like hearing. The top reply was blunt: clients do not care about internal numbers like ticket counts or first response time. They care whether their people can work without interruption, whether their data is secure, and whether projects land on time. Everything else is the MSP's homework.
That does not make the internal metrics useless. It means the client report translates them. First response time becomes hours of staff time spent waiting on IT this month. Patch compliance becomes "every laptop is on current patches, and the three exceptions are the CNC machine, the badge server and the CEO's old MacBook". Backup success rate becomes "restore test passed on the 14th, 22 minutes to a working file server". The vCIO reporting conversation is where those translated numbers belong.
Metrics That Get Gamed
Every metric with a target attached will be gamed, usually without anyone meaning to. The pattern is old enough to have a name, Goodhart's law: once a measure becomes a target, it stops being a good measure. The fix is never to drop the metric. It is to pair it with the number that exposes the shortcut.
Resolution time gets gamed by closing tickets fast. Reopen rate exposes it. SLA attainment gets gamed by parking tickets in "waiting on customer", where the clock stops. Time in pending exposes it. Backlog gets gamed by auto-closing anything without a reply in three days. Auto-close count exposes it. Utilization gets gamed by logging time generously, and a desk at 100% utilization has no hours left for documentation, which shows up six months later as a rising escalation rate.
An October 2024 r/msp thread on picking three support KPIs produced a good example of a metric that resists gaming: a "crush rate", a rolling seven-day count of open minus closed tickets. It is hard to fake, because closing tickets that come back moves it the wrong way within the same week.
Building the Scorecard
Run every candidate through four gates before it earns a tile. It has an owner, one person, not a team. It has a data source you can query today, not one you plan to build. It has a target, written down, with the date it was set. And it changes a decision: if the number goes red and nobody does anything different, it is a report, not a KPI.
A starter set that passes those gates for a small internal team: first response time by priority, resolution time by priority, reopen rate, backlog trend, availability for the two services people would notice, patch compliance, restore test date, MFA coverage. Eight tiles. For an MSP, per client: the same eight plus tickets per user per month and reactive hours per endpoint. Starting thresholds are just that, starting points: a reopen rate under 5%, patch compliance above 95% inside 14 days, a restore test inside the last 30 days. Move them once you have a quarter of your own data.
Coverage numbers are the ones that get counted by hand and drift. A script that checks agent state, patch level and encryption status across a client's devices and collects the output is how OpenFrame keeps them current without a spreadsheet. Whatever runs the script, the rule is the same: the scorecard reads from the system, never from a person's memory of the system.
Start With Three
If eight still feels like forty, start with three: first response time by priority, patch compliance with exceptions named, and the date of the last restore test. One tells you how the desk feels to users, one tells you how exposed the fleet is, and one tells you whether the backup you pay for would work. Add the rest a quarter at a time, and drop any tile that has not changed a decision in six months.
For the process that produces clean incident timestamps, read the incident management guide. For the promise you make from these numbers, read what an SLA is.
FAQ
What is the difference between an IT metric and a KPI?
A metric is any number you can count from a system, such as tickets opened or minutes to first reply. A KPI is a metric with an owner, a target and a decision attached: when it crosses the target, someone does something different. Every KPI is a metric; only a few metrics deserve to be KPIs.
How many IT KPIs should a team track?
Five to ten per function is the working rule. Past that, the dashboard stops changing behavior. A small internal IT team can run on eight: response and resolution time by priority, reopen rate, backlog trend, availability, patch compliance, restore test date and MFA coverage.
What is a good first response time for a help desk?
It depends on priority, which is why a single average is misleading. Common starting targets are 15 minutes for a P1 outage, one hour for a P2 that blocks one person's work, and four business hours for routine requests. Measure your own numbers for a quarter, then set targets you can hit nine times out of ten.
Which IT metrics should an MSP report to clients?
Translated ones. Clients care whether their people can work, whether their data is safe and whether projects land on time, so report hours lost waiting on IT, patch compliance with the exceptions named, the date and duration of the last restore test, and project milestones. Keep first response time, utilization and ticket counts for internal review.
Conrad Lunderstedt
Solution Architect
I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.
