A disk fills up on a Friday night, a scheduled job stops writing, and the first person to notice is a customer on Monday morning. Every one of those events left a trail somewhere, and something could have said so three days earlier. That gap between what a system already knows and what anyone gets told is the job of IT monitoring, and this guide covers what to watch across endpoints, network and apps, where to set the thresholds, and who gets woken up.
TL;DR
- IT monitoring. Collecting health and performance data from endpoints, network gear, applications and logs, then turning it into a dashboard, a ticket or a page depending on what it means.
- Four layers. Endpoints and servers, network, applications, logs. Each answers a different question, so skipping one leaves a blind spot.
- Thresholds. Start from a small published set, then tune to your own baseline. Alert on trend for disks, not just a fixed line.
- Routing. Page for things that need a human now, ticket for things that need a human today, dashboard for everything else.
- Rollout. Inventory, then thresholds, then routing, then a monthly review of what fired and what got ignored.
What Is IT Monitoring?
Google's SRE book defines monitoring as "collecting, processing, aggregating, and displaying real-time quantitative data about a system, such as query counts and types, error counts and types, processing times, and server lifetimes." Strip the Google scale out of that sentence and it still holds for a 40-person office. The data is already there. Monitoring is the part where someone collects it and decides what it means.
The same chapter draws the line that matters most for a small team. Monitoring has to answer two questions: what's broken, and why. The "what" is a symptom, the thing a user would notice, like a slow app or a login that fails. The "why" is a cause, like a database with a full transaction log. One layer's symptom is another layer's cause, which is why watching only servers or only applications leaves you guessing.
It also names two ways of looking. Black-box monitoring tests "externally visible behavior as a user would see it," a synthetic login or a ping to the office router. White-box monitoring reads "metrics exposed by the internals of the system," CPU counters, queue depths, event logs. You need both. Black-box tells you a customer is affected right now. White-box tells you what to fix and, on a good day, warns you before the black-box check goes red.
There's one more distinction worth settling early: IT monitoring is broader than infrastructure monitoring. Our guide to infrastructure monitoring goes deep on servers, storage and virtual hosts. This post is the map those pieces sit on, with endpoints, network, applications and logs treated as four layers of one system.
The Four Layers and What to Watch on Each
A layer is a place where a failure shows up first. The table below is the minimum set for a small IT team or an MSP running client fleets, and the sections after it explain the numbers.
| Layer | Watch first | What it tells you | Typical collector |
|---|---|---|---|
| Endpoints and servers | CPU, memory, disk free and disk latency, service state, agent heartbeat, pending reboot | A machine is short of a resource or about to stop doing its job | RMM agent, OS performance counters |
| Network | Interface errors and discards, link utilization, latency and packet loss to key targets, device reachability, DNS and DHCP health | Users can't reach things, or can reach them slowly | SNMP polling, flow data, synthetic pings |
| Applications | Request rate, error rate, response time, certificate expiry, queue depth, scheduled job success | The thing people use is failing or slow, whatever the servers say | Synthetic checks, APM, app logs |
| Logs and security events | Failed sign-ins, admin group changes, service installs, audit log cleared, backup job results | Something happened that no metric would show | Windows Event Forwarding, syslog, cloud audit logs |
Endpoints and Servers
Brendan Gregg's USE method is the shortest checklist for this layer: for every resource, check utilization, saturation and errors. Utilization is "the average time that the resource was busy servicing work." Saturation is "the degree to which the resource has extra work which it can't service, often queued." Errors are "the count of error events." Apply that to CPU, memory, disk and the network interface on each machine and you have covered the hardware.
The practical list for a Windows fleet is short. CPU sustained over a threshold for longer than a few minutes, not a spike. Available memory in absolute terms, because a percentage on a 4 GB laptop and a 256 GB host mean different things. Free disk space on the system drive, and disk latency on anything that runs a database. The state of services that should always be running, and the state of the monitoring agent itself, since an agent that stopped reporting looks exactly like a healthy machine.
Microsoft's own VM alert guidance is a useful sanity check because it names the metrics its recommended rules use: Percentage CPU, Available Memory Bytes, and Network In and Out totals from the host, plus a VM availability metric and a per-minute agent heartbeat. It also says those recommended rules "won't provide sufficient alerting for most enterprise implementations," because they cover the machine and not the workloads on it. That's the right mental model for any RMM's default monitor set too: a floor, not a finished design.
Two things belong on this layer that vendors rarely put in the default set. Pending reboots, because a server that has been waiting 40 days to finish a patch is a patch you don't have. And certificate and licence expiry dates on the machines that hold them, because the alert you want is 30 days out, not the morning the RDP gateway stops answering.
Network
Network monitoring has two halves, and small teams usually run only one of them. The first is device health: is the switch up, are its interfaces throwing errors or discards, how full is the uplink. That comes from SNMP, and our guide to SNMP OIDs covers which counters to poll and why v3 matters.
The second half is path health: can users reach the things they need, and how fast. A synthetic ping and a DNS lookup from inside each site to the file server, the identity provider and a public target every minute tells you more about the user's day than any switch counter. When the ISP degrades, the switch is fine and the users are not. Packet loss above a percent or two, and latency that doubles against its own baseline, are the signals to alert on.
Watch DHCP scope usage and DNS response time on the same schedule. A scope that runs out of addresses looks like a Wi-Fi outage to everyone in the building, and DNS that takes 800 ms turns every application into a slow application.
Applications and Services
For anything request-driven, the RED method is the checklist. Tom Wilkie's three signals are rate, "the number of requests per second," errors, "the number of those requests that are failing," and duration, "the amount of time those requests take." As he put it, USE cares about your machines and RED cares about your users. Google's four golden signals say the same with one addition: latency, traffic, errors and saturation, where saturation is how full the service is.
IBM's short explainer covers the four signals and why error rate beats CPU as the first thing to page on:
For a small team that doesn't run its own code, application monitoring mostly means synthetic checks. Log in to the line-of-business app every five minutes from a probe and time it. Hit the public website and check for a string on the page, not just a 200. Check that the nightly job wrote its output file, that the backup job reported success, and that the mail queue isn't growing. Each of these is a black-box test that catches what the server metrics miss, and each maps to a line in the client's service level agreement.
Logs and Security Events
Metrics tell you a machine is busy. Logs tell you what it did. Some of the most useful monitors on a Windows fleet aren't thresholds at all: an audit log cleared, a new member added to Domain Admins, a service installed on a server that shouldn't get one, a burst of failed sign-ins followed by a lockout. Those are single events that should raise a ticket or a page the moment they appear.
The catch is that a log only counts as monitored if it leaves the machine. Local event logs overwrite themselves, and an attacker clears them. Our guide to log management covers Windows Event Forwarding, syslog for network gear, and which event IDs to collect first, so this post won't repeat it. The one point to carry over: monitor the collectors too. A source that goes quiet looks identical to a quiet week.
Thresholds That Mean Something
A threshold is a guess about when a number becomes a problem. The default guesses in most tools are fine for the first week and wrong for years afterwards, because nobody revisits them. Three rules keep them useful.
First, measure against a baseline before you pick a line. Run the collectors for two weeks with no alerts, then look at what normal looks like per machine class. A file server at 90% memory is Windows caching files as designed. A domain controller at 90% memory is a problem. The same number, two meanings.
Second, alert on duration, not on a single sample. CPU at 100% for 30 seconds is a backup job. CPU at 90% for 15 minutes is a runaway process, and our guide on lowering CPU usage covers finding which one. Nearly every monitoring tool has a "for N minutes" or "consecutive breaches" option, and it removes more noise than any other single setting.
Third, for anything that fills up, alert on the trend and not just the level. Disk space is the classic case, and the r/sysadmin thread below is the whole debate in 40 comments. The opening question asks whether 15 GB free is too generous a floor. The top reply argues for 20 GB on both clients and servers because "disk is far cheaper than an outage." The next argues for 10% and 10 GB, whichever is worse, because a percentage alone misfires on giant disks. The comment that gets the most agreement points out that a 1 TB disk with 20 GB left and no growth needs nothing, while 900 GB free at 100 GB a day is urgent, so the alert you want is "30 days until 20 GB left."
The table below is a starting set for a Windows-heavy fleet, drawn from that thread and the metric set Microsoft's recommended rules use. Every row is a starting point to tune against your baseline, not a standard.
| Metric | Warning | Alert | Duration | Note |
|---|---|---|---|---|
| System disk free | Under 25% or 20 GB | Under 10% or 10 GB | Any | Add a trend rule: days until floor at current growth |
| CPU | Over 80% | Over 90% | 15 min | Exclude known backup and scan windows |
| Available memory | Under 15% | Under 1 GB absolute | 10 min | Absolute floor on servers, percent on laptops |
| Disk latency (database hosts) | Over 20 ms | Over 50 ms | 5 min | Read and write separately |
| Interface errors or discards | Any increase | Sustained increase | 5 min | Almost always a cable, an SFP or a duplex mismatch |
| Packet loss to key targets | Over 1% | Over 3% | 5 min | Per site, per target |
| Agent heartbeat | Missed 5 min | Missed 15 min | Any | Separate from "machine off" for laptops |
| Certificate expiry | 30 days | 14 days | Any | Include internal CA and RDP gateway certs |
| Backup job | Warning status | Failed or no report in 26 h | Any | "No report" is the one people forget |
Static Line vs Trend
The difference between the two disk rules is the difference between three days of warning and thirty. A static rule at 10% fires when the disk has already reached 10%, and on a server that grows 3% a week that's about three days before it's full. A trend rule that projects the current growth rate fires the moment the projection crosses "under 20 GB within 30 days," which on the same server is a month out. Same disk, same data, one alert lands on a Tuesday afternoon and the other lands at 02:00 on a Sunday.
Few RMM monitors do the projection natively, and the thread above says as much. The workaround is a scheduled script that reads free space daily, keeps a short history per machine, and raises a ticket when the slope crosses the line. That's a few lines of PowerShell and a place to store thirty numbers per disk. OpenFrame can run a script like that across a client's devices and collect each machine's output in one place, which is a fast way to see the fleet's growth rates side by side.
Alert Routing: Page, Ticket or Dashboard
Once a threshold trips, the next decision is who hears about it and how. Google's SRE book sets the bar for the loudest option: "every page should be actionable," "every page response should require intelligence," and pages should be about a novel problem. If an alert doesn't clear that bar, it isn't a page. It might still be a ticket, or a line on a dashboard, or nothing.
That gives you three destinations. A page interrupts a person now: the domain controller is down, the backup target is unreachable, the audit log was cleared. A ticket queues for a person today: disk trending toward the floor, a certificate at 30 days, a service that restarted twice overnight. A dashboard is for everything you want to see when you're already looking: utilization graphs, patch status, the count of machines that missed a heartbeat this week.
The failure mode is well documented in this r/sysadmin thread from April 2026. The poster describes a monitoring setup that generates so many alerts that the team ignores them, then asks whether to tune the alerts down or enforce stricter response. The top reply, at 38 upvotes, is blunt: if you're not taking action, you shouldn't be getting that alert, so start by removing everything informational-only. The second says the same in fewer words. The one dissent notes that some alerts need context before you know whether they're yours to act on, which is an argument for enrichment, not for more paging.
Routing has a second dimension: correlation. One switch going down takes 40 endpoints off the network, and a monitor that pages 41 times has told you nothing 40 times. Parent-child dependencies in the monitoring tool, where the switch is the parent of everything behind it, collapse that to one page. The same goes for a site's WAN link and everything at that site. Set the dependencies before you set the thresholds, because the dependencies decide how many alerts a single failure produces. Our guide to alert fatigue covers the tuning loop in detail.
Then decide who owns each route. A page goes to the on-call phone, and the rota has to be written down somewhere the tool can read it. A ticket goes to a queue with a priority attached, which is where the incident management process picks it up. A dashboard alert goes nowhere, and that's the point.
Building the Monitoring Dashboard
A monitoring dashboard has one job: let someone who just walked in see what's wrong in ten seconds. That rules out most default dashboards, which are a wall of gauges arranged by the tool's data model instead of by the question a person is asking.
Build two views. The first is the wall view, one screen, no scrolling: a status tile per site or per client, the count of open pages and tickets, and the three or four service checks that matter most (identity provider, mail, the main line-of-business app, internet at each site). Green, yellow, red, and a timestamp so nobody trusts a stale screen. The second is the triage view, which opens when a tile goes red: the machine or site, its last hour of CPU, memory, disk and network, the last 20 events, and the ticket link.
Put the four golden signals on the application tiles, not the server tiles. Rate, errors and latency for the app tell you about users. CPU on the app server tells you about the server. Both matter, but the tile a manager looks at should show the first one.
For MSPs the dashboard also has to answer "which client" before "which machine." A per-client roll-up with a drill-down beats one global list every time, and it's the view the client can be shown at the quarterly review without exposing anything sensitive.
Monitoring vs Observability
The two words get used as if one replaced the other. They didn't. Monitoring asks known questions of a system: is the disk under 10%, did the job run, is the site up. Observability is the property of a system that lets you ask questions you didn't plan for, usually by keeping enough metrics, logs and traces that a new question can be answered from data you already have.
Better Stack's four-minute explainer draws the line cleanly:
For a small IT team the practical reading is this. Get monitoring right first: the four layers, sane thresholds, routing that people trust. Observability then comes almost for free from keeping the data longer and searchable, which is the log management work already described. Buying an observability platform before the monitoring basics are in place gets you a bigger dashboard with the same blind spots.
Choosing the Stack Without Buying Four Tools
Nobody needs a separate product per layer, but nobody has one product that does all four well either. The realistic shape for a small team or an MSP is two or three tools with clear ownership.
The RMM covers endpoints and servers: agents, performance counters, service checks, scripts and patch state. Our explainer on what RMM does covers where its monitoring stops.
A network monitor covers the gear: SNMP polling, interface counters, topology, and synthetic path checks from each site. The network management software roundup compares the options with pricing.
Cloud workloads usually come with their platform's own monitoring. Our guide to cloud monitoring tools covers getting Azure, AWS and SaaS signals into one place instead of three consoles.
That leaves applications and logs. Synthetic checks are cheap and many network monitors include them. Logs need a collector and a place to search, and the log management guide above covers open-source and commercial options. The test for the whole stack is the wall dashboard: if you can't build the ten-second view from the tools you have, the gap is in the tools, not the dashboard.
Two outage figures put the layers in proportion. Uptime Institute's 2025 Annual Outage Analysis (May 2025) found that nearly 40% of organizations had suffered a major outage caused by human error in the past three years, and that IT and networking issues accounted for 23% of impactful outages in 2024. Monitoring doesn't stop the human error. It shortens the time between the mistake and someone noticing.
A 30-Day Rollout Checklist
Monitoring projects fail by starting with thresholds. Start with the inventory, and let the thresholds come last.
| Week | Do | Done when |
|---|---|---|
| 1 | Inventory every endpoint, server, network device and cloud workload; install or verify agents; confirm SNMP access to gear; list the five services users depend on most | Every asset reports a heartbeat and the five services have a synthetic check |
| 2 | Collect with no alerts; record per-class baselines for CPU, memory, disk growth, latency and packet loss; set parent-child dependencies for switches and WAN links | A baseline sheet exists per machine class and per site |
| 3 | Set thresholds from the baselines with durations; add trend rules for disks; add event monitors for audit log cleared, admin group changes and backup results; define page, ticket and dashboard routes and the on-call rota | Every alert has a destination and an owner |
| 4 | Build the wall and triage dashboards; run a tabletop on one page and one ticket end to end; review every alert that fired in week 3 and delete or demote the ones nobody acted on | The alert list is shorter than it was on day 15 |
Then repeat week 4 every month. The review question is the same each time: which alerts fired, which got actioned, and which got ignored. Ignored alerts get demoted to a ticket or a dashboard line, or removed. New machines and new services get the same treatment as the first batch, or the noise creeps back within a quarter.
Where to Go Next
IT monitoring is four layers, a small set of thresholds tuned to your own baseline, and a routing rule that keeps pages rare. Get the inventory and the dependencies right first, then the thresholds, then the dashboards. If the server side is where your gaps are, the infrastructure monitoring guide is the deep dive. If the alerts are the problem, start with the alert fatigue guide and cut the list before you add anything to it.

Aliaska Varieva
Head of Platform
Hi! I’m Aliaska, and I’ve been working as a software engineer (mostly Java + a bit Kotlin) for over 8 years now. I mostly spend my time building backend services, integrating systems, fixing bugs (the fun part 🙃), and making sure things don’t fall apart behind the scenes.
