Flamingo Raises $4.5M Seed Round

Skip to content

A disk fills up on a Friday night, a scheduled job stops writing, and the first person to notice is a customer on Monday morning. Every one of those events left a trail somewhere, and something could have said so three days earlier. That gap between what a system already knows and what anyone gets told is the job of IT monitoring, and this guide covers what to watch across endpoints, network and apps, where to set the thresholds, and who gets woken up.

TL;DR

  • IT monitoring. Collecting health and performance data from endpoints, network gear, applications and logs, then turning it into a dashboard, a ticket or a page depending on what it means.
  • Four layers. Endpoints and servers, network, applications, logs. Each answers a different question, so skipping one leaves a blind spot.
  • Thresholds. Start from a small published set, then tune to your own baseline. Alert on trend for disks, not just a fixed line.
  • Routing. Page for things that need a human now, ticket for things that need a human today, dashboard for everything else.
  • Rollout. Inventory, then thresholds, then routing, then a monthly review of what fired and what got ignored.

What Is IT Monitoring?

Google's SRE book defines monitoring as "collecting, processing, aggregating, and displaying real-time quantitative data about a system, such as query counts and types, error counts and types, processing times, and server lifetimes." Strip the Google scale out of that sentence and it still holds for a 40-person office. The data is already there. Monitoring is the part where someone collects it and decides what it means.

The same chapter draws the line that matters most for a small team. Monitoring has to answer two questions: what's broken, and why. The "what" is a symptom, the thing a user would notice, like a slow app or a login that fails. The "why" is a cause, like a database with a full transaction log. One layer's symptom is another layer's cause, which is why watching only servers or only applications leaves you guessing.

It also names two ways of looking. Black-box monitoring tests "externally visible behavior as a user would see it," a synthetic login or a ping to the office router. White-box monitoring reads "metrics exposed by the internals of the system," CPU counters, queue depths, event logs. You need both. Black-box tells you a customer is affected right now. White-box tells you what to fix and, on a good day, warns you before the black-box check goes red.

There's one more distinction worth settling early: IT monitoring is broader than infrastructure monitoring. Our guide to infrastructure monitoring goes deep on servers, storage and virtual hosts. This post is the map those pieces sit on, with endpoints, network, applications and logs treated as four layers of one system.

The Four Layers and What to Watch on Each

A layer is a place where a failure shows up first. The table below is the minimum set for a small IT team or an MSP running client fleets, and the sections after it explain the numbers.

LayerWatch firstWhat it tells youTypical collector
Endpoints and serversCPU, memory, disk free and disk latency, service state, agent heartbeat, pending rebootA machine is short of a resource or about to stop doing its jobRMM agent, OS performance counters
NetworkInterface errors and discards, link utilization, latency and packet loss to key targets, device reachability, DNS and DHCP healthUsers can't reach things, or can reach them slowlySNMP polling, flow data, synthetic pings
ApplicationsRequest rate, error rate, response time, certificate expiry, queue depth, scheduled job successThe thing people use is failing or slow, whatever the servers saySynthetic checks, APM, app logs
Logs and security eventsFailed sign-ins, admin group changes, service installs, audit log cleared, backup job resultsSomething happened that no metric would showWindows Event Forwarding, syslog, cloud audit logs

Endpoints and Servers

Brendan Gregg's USE method is the shortest checklist for this layer: for every resource, check utilization, saturation and errors. Utilization is "the average time that the resource was busy servicing work." Saturation is "the degree to which the resource has extra work which it can't service, often queued." Errors are "the count of error events." Apply that to CPU, memory, disk and the network interface on each machine and you have covered the hardware.

The practical list for a Windows fleet is short. CPU sustained over a threshold for longer than a few minutes, not a spike. Available memory in absolute terms, because a percentage on a 4 GB laptop and a 256 GB host mean different things. Free disk space on the system drive, and disk latency on anything that runs a database. The state of services that should always be running, and the state of the monitoring agent itself, since an agent that stopped reporting looks exactly like a healthy machine.

Microsoft's own VM alert guidance is a useful sanity check because it names the metrics its recommended rules use: Percentage CPU, Available Memory Bytes, and Network In and Out totals from the host, plus a VM availability metric and a per-minute agent heartbeat. It also says those recommended rules "won't provide sufficient alerting for most enterprise implementations," because they cover the machine and not the workloads on it. That's the right mental model for any RMM's default monitor set too: a floor, not a finished design.

Two things belong on this layer that vendors rarely put in the default set. Pending reboots, because a server that has been waiting 40 days to finish a patch is a patch you don't have. And certificate and licence expiry dates on the machines that hold them, because the alert you want is 30 days out, not the morning the RDP gateway stops answering.

Network

Network monitoring has two halves, and small teams usually run only one of them. The first is device health: is the switch up, are its interfaces throwing errors or discards, how full is the uplink. That comes from SNMP, and our guide to SNMP OIDs covers which counters to poll and why v3 matters.

The second half is path health: can users reach the things they need, and how fast. A synthetic ping and a DNS lookup from inside each site to the file server, the identity provider and a public target every minute tells you more about the user's day than any switch counter. When the ISP degrades, the switch is fine and the users are not. Packet loss above a percent or two, and latency that doubles against its own baseline, are the signals to alert on.

Watch DHCP scope usage and DNS response time on the same schedule. A scope that runs out of addresses looks like a Wi-Fi outage to everyone in the building, and DNS that takes 800 ms turns every application into a slow application.

Applications and Services

For anything request-driven, the RED method is the checklist. Tom Wilkie's three signals are rate, "the number of requests per second," errors, "the number of those requests that are failing," and duration, "the amount of time those requests take." As he put it, USE cares about your machines and RED cares about your users. Google's four golden signals say the same with one addition: latency, traffic, errors and saturation, where saturation is how full the service is.

IBM's short explainer covers the four signals and why error rate beats CPU as the first thing to page on:

For a small team that doesn't run its own code, application monitoring mostly means synthetic checks. Log in to the line-of-business app every five minutes from a probe and time it. Hit the public website and check for a string on the page, not just a 200. Check that the nightly job wrote its output file, that the backup job reported success, and that the mail queue isn't growing. Each of these is a black-box test that catches what the server metrics miss, and each maps to a line in the client's service level agreement.

Logs and Security Events

Metrics tell you a machine is busy. Logs tell you what it did. Some of the most useful monitors on a Windows fleet aren't thresholds at all: an audit log cleared, a new member added to Domain Admins, a service installed on a server that shouldn't get one, a burst of failed sign-ins followed by a lockout. Those are single events that should raise a ticket or a page the moment they appear.

The catch is that a log only counts as monitored if it leaves the machine. Local event logs overwrite themselves, and an attacker clears them. Our guide to log management covers Windows Event Forwarding, syslog for network gear, and which event IDs to collect first, so this post won't repeat it. The one point to carry over: monitor the collectors too. A source that goes quiet looks identical to a quiet week.

Thresholds That Mean Something

A threshold is a guess about when a number becomes a problem. The default guesses in most tools are fine for the first week and wrong for years afterwards, because nobody revisits them. Three rules keep them useful.

First, measure against a baseline before you pick a line. Run the collectors for two weeks with no alerts, then look at what normal looks like per machine class. A file server at 90% memory is Windows caching files as designed. A domain controller at 90% memory is a problem. The same number, two meanings.

Second, alert on duration, not on a single sample. CPU at 100% for 30 seconds is a backup job. CPU at 90% for 15 minutes is a runaway process, and our guide on lowering CPU usage covers finding which one. Nearly every monitoring tool has a "for N minutes" or "consecutive breaches" option, and it removes more noise than any other single setting.

Third, for anything that fills up, alert on the trend and not just the level. Disk space is the classic case, and the r/sysadmin thread below is the whole debate in 40 comments. The opening question asks whether 15 GB free is too generous a floor. The top reply argues for 20 GB on both clients and servers because "disk is far cheaper than an outage." The next argues for 10% and 10 GB, whichever is worse, because a percentage alone misfires on giant disks. The comment that gets the most agreement points out that a 1 TB disk with 20 GB left and no growth needs nothing, while 900 GB free at 100 GB a day is urgent, so the alert you want is "30 days until 20 GB left."

The table below is a starting set for a Windows-heavy fleet, drawn from that thread and the metric set Microsoft's recommended rules use. Every row is a starting point to tune against your baseline, not a standard.

MetricWarningAlertDurationNote
System disk freeUnder 25% or 20 GBUnder 10% or 10 GBAnyAdd a trend rule: days until floor at current growth
CPUOver 80%Over 90%15 minExclude known backup and scan windows
Available memoryUnder 15%Under 1 GB absolute10 minAbsolute floor on servers, percent on laptops
Disk latency (database hosts)Over 20 msOver 50 ms5 minRead and write separately
Interface errors or discardsAny increaseSustained increase5 minAlmost always a cable, an SFP or a duplex mismatch
Packet loss to key targetsOver 1%Over 3%5 minPer site, per target
Agent heartbeatMissed 5 minMissed 15 minAnySeparate from "machine off" for laptops
Certificate expiry30 days14 daysAnyInclude internal CA and RDP gateway certs
Backup jobWarning statusFailed or no report in 26 hAny"No report" is the one people forget

Static Line vs Trend

The difference between the two disk rules is the difference between three days of warning and thirty. A static rule at 10% fires when the disk has already reached 10%, and on a server that grows 3% a week that's about three days before it's full. A trend rule that projects the current growth rate fires the moment the projection crosses "under 20 GB within 30 days," which on the same server is a month out. Same disk, same data, one alert lands on a Tuesday afternoon and the other lands at 02:00 on a Sunday.

Few RMM monitors do the projection natively, and the thread above says as much. The workaround is a scheduled script that reads free space daily, keeps a short history per machine, and raises a ticket when the slope crosses the line. That's a few lines of PowerShell and a place to store thirty numbers per disk. OpenFrame can run a script like that across a client's devices and collect each machine's output in one place, which is a fast way to see the fleet's growth rates side by side.

Alert Routing: Page, Ticket or Dashboard

Once a threshold trips, the next decision is who hears about it and how. Google's SRE book sets the bar for the loudest option: "every page should be actionable," "every page response should require intelligence," and pages should be about a novel problem. If an alert doesn't clear that bar, it isn't a page. It might still be a ticket, or a line on a dashboard, or nothing.

That gives you three destinations. A page interrupts a person now: the domain controller is down, the backup target is unreachable, the audit log was cleared. A ticket queues for a person today: disk trending toward the floor, a certificate at 30 days, a service that restarted twice overnight. A dashboard is for everything you want to see when you're already looking: utilization graphs, patch status, the count of machines that missed a heartbeat this week.

The failure mode is well documented in this r/sysadmin thread from April 2026. The poster describes a monitoring setup that generates so many alerts that the team ignores them, then asks whether to tune the alerts down or enforce stricter response. The top reply, at 38 upvotes, is blunt: if you're not taking action, you shouldn't be getting that alert, so start by removing everything informational-only. The second says the same in fewer words. The one dissent notes that some alerts need context before you know whether they're yours to act on, which is an argument for enrichment, not for more paging.

Routing has a second dimension: correlation. One switch going down takes 40 endpoints off the network, and a monitor that pages 41 times has told you nothing 40 times. Parent-child dependencies in the monitoring tool, where the switch is the parent of everything behind it, collapse that to one page. The same goes for a site's WAN link and everything at that site. Set the dependencies before you set the thresholds, because the dependencies decide how many alerts a single failure produces. Our guide to alert fatigue covers the tuning loop in detail.

Then decide who owns each route. A page goes to the on-call phone, and the rota has to be written down somewhere the tool can read it. A ticket goes to a queue with a priority attached, which is where the incident management process picks it up. A dashboard alert goes nowhere, and that's the point.

Building the Monitoring Dashboard

A monitoring dashboard has one job: let someone who just walked in see what's wrong in ten seconds. That rules out most default dashboards, which are a wall of gauges arranged by the tool's data model instead of by the question a person is asking.

Build two views. The first is the wall view, one screen, no scrolling: a status tile per site or per client, the count of open pages and tickets, and the three or four service checks that matter most (identity provider, mail, the main line-of-business app, internet at each site). Green, yellow, red, and a timestamp so nobody trusts a stale screen. The second is the triage view, which opens when a tile goes red: the machine or site, its last hour of CPU, memory, disk and network, the last 20 events, and the ticket link.

Put the four golden signals on the application tiles, not the server tiles. Rate, errors and latency for the app tell you about users. CPU on the app server tells you about the server. Both matter, but the tile a manager looks at should show the first one.

For MSPs the dashboard also has to answer "which client" before "which machine." A per-client roll-up with a drill-down beats one global list every time, and it's the view the client can be shown at the quarterly review without exposing anything sensitive.

Monitoring vs Observability

The two words get used as if one replaced the other. They didn't. Monitoring asks known questions of a system: is the disk under 10%, did the job run, is the site up. Observability is the property of a system that lets you ask questions you didn't plan for, usually by keeping enough metrics, logs and traces that a new question can be answered from data you already have.

Better Stack's four-minute explainer draws the line cleanly:

For a small IT team the practical reading is this. Get monitoring right first: the four layers, sane thresholds, routing that people trust. Observability then comes almost for free from keeping the data longer and searchable, which is the log management work already described. Buying an observability platform before the monitoring basics are in place gets you a bigger dashboard with the same blind spots.

Choosing the Stack Without Buying Four Tools

Nobody needs a separate product per layer, but nobody has one product that does all four well either. The realistic shape for a small team or an MSP is two or three tools with clear ownership.

The RMM covers endpoints and servers: agents, performance counters, service checks, scripts and patch state. Our explainer on what RMM does covers where its monitoring stops.

A network monitor covers the gear: SNMP polling, interface counters, topology, and synthetic path checks from each site. The network management software roundup compares the options with pricing.

Cloud workloads usually come with their platform's own monitoring. Our guide to cloud monitoring tools covers getting Azure, AWS and SaaS signals into one place instead of three consoles.

That leaves applications and logs. Synthetic checks are cheap and many network monitors include them. Logs need a collector and a place to search, and the log management guide above covers open-source and commercial options. The test for the whole stack is the wall dashboard: if you can't build the ten-second view from the tools you have, the gap is in the tools, not the dashboard.

Two outage figures put the layers in proportion. Uptime Institute's 2025 Annual Outage Analysis (May 2025) found that nearly 40% of organizations had suffered a major outage caused by human error in the past three years, and that IT and networking issues accounted for 23% of impactful outages in 2024. Monitoring doesn't stop the human error. It shortens the time between the mistake and someone noticing.

A 30-Day Rollout Checklist

Monitoring projects fail by starting with thresholds. Start with the inventory, and let the thresholds come last.

WeekDoDone when
1Inventory every endpoint, server, network device and cloud workload; install or verify agents; confirm SNMP access to gear; list the five services users depend on mostEvery asset reports a heartbeat and the five services have a synthetic check
2Collect with no alerts; record per-class baselines for CPU, memory, disk growth, latency and packet loss; set parent-child dependencies for switches and WAN linksA baseline sheet exists per machine class and per site
3Set thresholds from the baselines with durations; add trend rules for disks; add event monitors for audit log cleared, admin group changes and backup results; define page, ticket and dashboard routes and the on-call rotaEvery alert has a destination and an owner
4Build the wall and triage dashboards; run a tabletop on one page and one ticket end to end; review every alert that fired in week 3 and delete or demote the ones nobody acted onThe alert list is shorter than it was on day 15

Then repeat week 4 every month. The review question is the same each time: which alerts fired, which got actioned, and which got ignored. Ignored alerts get demoted to a ticket or a dashboard line, or removed. New machines and new services get the same treatment as the first batch, or the noise creeps back within a quarter.

Where to Go Next

IT monitoring is four layers, a small set of thresholds tuned to your own baseline, and a routing rule that keeps pages rare. Get the inventory and the dependencies right first, then the thresholds, then the dashboards. If the server side is where your gaps are, the infrastructure monitoring guide is the deep dive. If the alerts are the problem, start with the alert fatigue guide and cut the list before you add anything to it.

Aliaska Varieva

Aliaska Varieva

Head of Platform

Hi! I’m Aliaska, and I’ve been working as a software engineer (mostly Java + a bit Kotlin) for over 8 years now. I mostly spend my time building backend services, integrating systems, fixing bugs (the fun part 🙃), and making sure things don’t fall apart behind the scenes.

Related Content

Blog Posts

Product Releases

Podcasts

Webinars

Case Studies

Events

Onboarding Guides

Frequently Asked Questions

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.

MSP AI Agents

On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.
Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.
IT monitoring is the practice of collecting health and performance data from endpoints, servers, network devices, applications and logs, then turning it into something a person can act on: a dashboard line, a ticket, or a page. Google's SRE book describes monitoring as collecting, processing, aggregating and displaying real-time quantitative data about a system. In a small IT team or an MSP, the data already exists; monitoring is deciding what it means and who hears about it.
Infrastructure monitoring covers servers, storage, virtual hosts and the network hardware they run on. IT monitoring is the wider map: it adds endpoints, applications and services as users experience them, and the logs and security events that no metric shows. Infrastructure monitoring is one layer of IT monitoring, usually the one with the most tooling already in place.
Start with a heartbeat from every device, free disk space on system drives, the state of services that must always run, backup job results, and a synthetic check on the three to five services users depend on most (identity, mail, the main line-of-business app, internet at each site). Add single-event monitors for an audit log being cleared and changes to admin groups. Expand from there once those alerts are trusted.
Collect for two weeks with no alerts and record what normal looks like per machine class and per site. Then set thresholds against that baseline, always with a duration so a 30-second spike does not fire. For anything that fills up, such as disks, add a trend rule that projects days until the floor rather than only a fixed percentage. Review every alert monthly and remove or demote the ones nobody acted on.
Only when a person needs to act now and the alert is actionable and novel: a domain controller down, a backup target unreachable, an audit log cleared. Anything that needs a person today becomes a ticket with a priority, and anything you only want to see when you are already looking belongs on a dashboard. Parent-child dependencies, such as a switch as the parent of the endpoints behind it, keep one failure from paging dozens of times.
Monitoring asks known questions of a system: is the disk under 10 percent, did the job run, is the site up. Observability is the property of a system that lets you answer questions you did not plan for, which usually means keeping enough metrics, logs and traces that a new question can be answered from data you already have. Get the monitoring basics right first; observability then comes mostly from retaining and searching the same data.