Flamingo Raises $4.5M Seed Round

Skip to content

Updated: October 2026

Every stack you run has a moment when something slows down, fills up or stops, and someone finds out. The only question is whether that someone is your monitoring or your user. This guide to IT infrastructure monitoring covers what to watch at each layer, how the data gets collected, which alerts are worth waking up for and which tools fit which job.

TL;DR

  • Definition. IT infrastructure monitoring is the continuous collection of health and performance data from servers, network devices, storage, cloud services and endpoints, with alerts when something needs a person.
  • Collection. Agents give the deepest data; SNMP, WMI, APIs and syslog cover what can't run one.
  • Metrics. Use utilization, saturation and errors for resources, and latency, traffic, errors and saturation for services.
  • Alerts. Page only on symptoms someone can act on. Everything else goes to a ticket or a dashboard.
  • Tools. Match the category to the job: RMM for endpoints, NMS for the network, APM for applications.

What IT Infrastructure Monitoring Covers

Infrastructure is everything an application needs to run that isn't the application itself. For a typical IT team that means six layers.

Servers come first: physical hosts, hypervisors and the virtual machines on them. You watch CPU, memory, disk, services and the event log. Network devices are next: switches, routers, firewalls, wireless controllers and the links between them. Here you care about interface errors, bandwidth, latency and whether the device answers at all.

Storage is its own layer, because a SAN or NAS can fail while every server looks fine. Cloud and SaaS form the fourth layer, from Azure and AWS resources to Microsoft 365 service health. Endpoints are the fifth: the laptops and desktops your users work on, which fail in smaller ways but far more often. The sixth layer is the services running on top: DNS, DHCP, Active Directory, backups, certificates and the line-of-business apps people call you about.

Monitoring all six is what separates infrastructure monitoring from a single-purpose tool. A network monitor tells you the switch port is fine. It can't tell you the file server behind it ran out of disk an hour ago. If you want the broader discipline around these layers, from procurement to lifecycle, our guide to IT infrastructure management covers it.

Agent vs Agentless: How the Data Gets Collected

Every monitoring tool has to get data off a device somehow. There are two families of methods, and a real environment uses both.

An agent is a small program installed on the device. It reads performance counters, logs, services and installed software locally, then pushes the results to a central server. Agents give you the deepest data and keep working when the device roams off the office network. The cost is deployment and upkeep: every agent is software you have to install, update and trust.

Agentless monitoring asks the device over the network instead. The main protocols are:

MethodWhat it readsWhere it fitsWatch out for
SNMP pollingCounters in a device's MIB, polled on UDP 161Switches, routers, firewalls, UPS units, printersv1 and v2c send the community string in cleartext
SNMP trapsEvents the device pushes on UDP 162Link down, fan failure, power eventsTraps are fire-and-forget; a lost trap is a lost event
WMI / WinRMWindows performance and configuration dataWindows servers without an agentNeeds credentials and firewall rules on every target
APIsMetrics from cloud platforms, hypervisors, SaaSAzure, AWS, vCenter, Microsoft 365Rate limits and token expiry
SyslogEvent messages sent to a collectorNetwork gear, Linux hosts, appliancesVolume; needs parsing to be useful
Flow dataWho talked to whom, how much (NetFlow, sFlow, IPFIX)Bandwidth questions, unusual trafficSampling hides small flows

SNMP deserves a warning. CISA's alert TA17-156A says that with SNMPv1 or v2, an adversary can sniff network traffic to learn the community string. Its recommendation is blunt: SNMPv3 should be the only version in use, because it can authenticate and encrypt. Treat any device still answering to "public" as a finding, not a setting.

WMI has its own quirks. Microsoft's WMI documentation notes that remote WMI connections run over DCOM, with WinRM as the alternative that uses the WS-Management protocol. DCOM needs dynamic ports, which is why agentless Windows monitoring so often dies at the firewall.

The practical rule: put agents on anything you own and can install software on, and use agentless methods for everything else. Network gear, appliances and cloud services will never run your agent.

Nagios' short explainer shows how SNMP managers, agents and MIBs fit together, if the protocol is new to anyone on the team:

The Metrics That Matter per Layer

Collecting everything is easy. Knowing which numbers mean trouble is the work. Two frameworks do most of the thinking for you.

For resources, use Brendan Gregg's USE method: "For every resource, check utilization, saturation, and errors." Utilization is how busy the resource is. Saturation is work queued that the resource can't serve yet. Errors are failures. A CPU at 90% with nothing queued is busy. A CPU at 70% with a long run queue is saturated, and that's the one users feel.

For services, use the four golden signals from Google's SRE book chapter on monitoring distributed systems: latency, traffic, errors and saturation. The chapter adds a detail worth keeping: track the latency of failed requests separately, because a slow error is worse than a fast one.

Applied to each layer, that gives you a short list:

LayerWatchWhy it mattersExample alert condition
ServerCPU run queue, memory pressure, disk free and growth rateSaturation shows up before utilization maxes outDisk projected to fill within days, not "disk at 90%"
NetworkInterface errors and discards, utilization, latency, device reachabilityErrors point at cabling, duplex or hardwareError count rising on an uplink over 15 minutes
StorageLatency per volume, free capacity, RAID or pool healthLatency hurts every VM on the arrayRead or write latency above your baseline for a sustained window
Cloud and SaaSResource health, quotas, service health noticesYou can't see the hardware, so you watch the provider's signalsA service health incident on a tenant you manage
EndpointsDisk space, patch status, security agent health, batteryMany small failures that become ticketsProtection disabled, or a device silent for days
ServicesResponse time, error rate, certificate expiry, backup job resultThis is what users experienceBackup job didn't report success last night

The example conditions are starting points, not standards. The right number for your environment comes from its baseline.

Disk space is the classic case for thinking in rates. A 2 TB volume at 90% has 200 GB left. If it grows 5 GB a day, that's six weeks and a scheduled job. If it grows 100 GB a day, that's two days and an emergency. The percentage is the same. Alert on the time to full instead.

IBM Technology's explainer walks through the golden signals with examples, and it's a good ten minutes for anyone setting thresholds:

Baselines: Know What Normal Looks Like

A threshold without a baseline is a guess. The same 80% CPU is normal for a build server at 2 a.m. and alarming for a domain controller at noon.

Collect two to four weeks of data before you set most thresholds. You want to see a full cycle of business days, weekends, month-end jobs and backup windows. Then set alerts relative to what you saw: a sustained deviation from the usual range, not a fixed number copied from a vendor template.

Add duration to every performance alert. A single sample above the line is noise. The same value held for 10 or 15 minutes is a trend. Duration windows are the cheapest way to cut false positives, and nearly every monitoring tool supports them.

Some tools offer dynamic thresholds that learn the baseline for you. They work well for metrics with a clear daily rhythm, like traffic or logins. They work badly for metrics that should never change, like free space on a volume that only grows. For those, static rules on the rate of change still win.

The stakes justify the effort. In Uptime Institute's 2025 annual survey, 57% of respondents said their most recent major outage cost more than $100,000, and for the second year running, 1 in 5 reported costs above $1 million. Catching a trend a day early is cheap by comparison.

Alert Design People Don't Ignore

Alert fatigue is how good monitoring dies. The tool works, the alerts fire, and nobody reads them anymore.

This r/sysadmin thread from April 2026 asks whether to tune alerts down or enforce stricter response times. The top reply settles it in one line: if you're not taking action, you shouldn't be getting that alert.

Google's SRE book puts the same rule in writing: "Every page should be actionable." It also says to spend more effort catching symptoms than causes. A symptom is what users feel: the website is slow, logins fail, the file share is gone. A cause is one of many reasons: high CPU, a full log volume, a flapping link. Page on symptoms. Keep causes on the dashboard and in the ticket, where they help the person investigating.

A workable scheme has three tiers:

TierGoes toResponseExamples
PagePhone of whoever is on callNow, any hourSite down, VPN unreachable, backup server offline
TicketThe helpdesk queueWithin business hoursDisk projected full this week, certificate expiring in 21 days
RecordDashboard and reportsReviewed weeklyCPU spikes that cleared, informational events

Three habits keep the tiers clean. Deduplicate, so one failed switch creates one alert instead of forty from the devices behind it. Auto-close tickets when the condition clears, so the queue reflects what's broken now. And review every page after the fact: if nobody had to do anything, demote the rule. Our guide to alert fatigue goes deeper on tuning RMM and SIEM noise.

Monitor What Stops Happening

Threshold alerts catch things that go wrong loudly. The failures that hurt most are quiet: a job that stopped running, a log that stopped arriving, a backup that last succeeded three weeks ago.

This August 2026 r/sysadmin thread starts with exactly that story. A nightly backup cron job had been failing or not running for a stretch. Uptime checks were green, nobody got a ticket, and the team found out when they needed a restore.

The fix is a dead man's switch, also called a heartbeat check. The job reports in when it finishes, and the monitor alerts when the report doesn't arrive on time. The replies in that thread cover the options, from hosted services to self-hosted push monitors.

Apply the same idea wherever silence means failure:

  • Backups. Alert on the absence of a success record, not only on a failure message.
  • Log sources. Alert when a firewall or domain controller stops sending logs. A quiet source is either broken or tampered with.
  • Scheduled tasks. Patch runs, sync jobs and report exports should all check in.
  • Agents. Flag devices that haven't reported in for days. They're either offline, reimaged or broken.
  • Certificates and licenses. Expiry dates are known in advance, so a ticket should open weeks before.

A Worked Example: One Office, One Bad Morning

Here's how the pieces fit together in an illustrative case. A 40-person office runs two Windows servers on one host, a firewall, two switches, Microsoft 365 and a cloud backup service.

At 08:40, users report that the shared drive is slow. Without monitoring, a tech starts at the laptop and works backwards. With it, the story is already on screen.

The storage layer shows read latency on the host's datastore climbing since 07:15. The server layer shows the file server's disk queue saturated, while CPU sits at 35%. The network layer is clean: no interface errors, normal bandwidth. So it isn't the switch, and it isn't the laptops.

The alert that paged the on-call tech was a symptom check: file share response time above its baseline for 15 minutes. The cause sat one layer down, on the dashboard. Last night's backup job ran long and was still reading from the datastore when the working day started.

Three changes come out of the review. The backup window moves earlier. A heartbeat check now opens a ticket if the job hasn't finished by 07:00. And the datastore latency rule becomes a ticket-tier alert, because it predicted the problem 85 minutes before anyone felt it.

Nothing in that story needed an expensive tool. It needed data from every layer in one timeline, and alerts sorted by who has to act.

Monitoring vs Observability

You'll see both words in vendor pitches, often used as synonyms. They overlap, but they answer different questions.

Monitoring asks known questions on a schedule. Is the server up? Is the disk filling? Is latency above the line? You decide in advance what to check, and the tool tells you when the answer changes.

Observability is about questions you didn't know you'd need to ask. It relies on rich telemetry, usually metrics, logs and traces, detailed enough to explore after something strange happens. It matters most for custom applications with many moving parts.

For infrastructure, monitoring does the heavy lifting. Servers, switches and backups fail in known ways, and known checks catch them. Add observability tooling when you run software complex enough that the known checks keep missing the cause.

Dashboards vs Alerts

Dashboards and alerts answer different questions, and mixing them up wastes both.

An alert answers: does someone need to act? It should be rare, specific and routed to a person. A dashboard answers: what's going on? It should be broad, visual and available when someone wants to look.

A dashboard nobody opens isn't monitoring. A wall screen in the office is useful during an incident and background noise the rest of the time. Build dashboards around questions people ask: "Is it the network or the server?", "Which client sites are degraded?", "What changed in the last hour?" Each panel should help answer one of them.

Put capacity trends on a separate dashboard and review it monthly. That's where disk growth, license counts and aging hardware show up, and where next quarter's budget request comes from.

Cloud and SaaS in the Same View

Cloud resources change how monitoring works. You can't install an agent on a managed database or a SaaS tenant. You read what the provider exposes: metrics APIs, activity logs and service health notices.

Pull those signals into the same place as your on-premises data where you can. An outage that starts in a cloud region and ends as a help desk flood is one incident, not two. Watch the provider's service health for every tenant you manage, because a Microsoft 365 incident explains a lot of "email is broken" tickets before anyone starts troubleshooting a laptop.

Cloud monitoring has its own tools and costs, and deserves its own planning. We're covering cloud monitoring tools in a separate guide. For infrastructure monitoring, the rule is simpler: whatever you can't see in your main tool, you need at least an alert rule for.

Tool Categories: Which Type Fits

No single category covers every layer equally well. Expect to run two or three tools, and put the effort into picking the combination.

CategoryStrongest atWeaker atExamples
RMMEndpoints and servers you manage, patching, scriptingDeep network and application metricsNinjaOne, Datto RMM, N-able
Network monitoring (NMS)SNMP devices, topology, bandwidth, flowEndpoints, application behaviourPRTG, LibreNMS, Auvik
Infrastructure and server monitoringHosts, services, custom checks at scaleEndpoint managementZabbix, Checkmk, Nagios, Icinga
APM and observabilityApplication latency, traces, logsHardware and network devicesDatadog, Dynatrace, New Relic
Open-source metrics stackFlexible time-series, custom dashboardsSetup and upkeep effortPrometheus with Grafana
Cloud-native monitoringDeep data about one providerAnything outside that providerAzure Monitor, Amazon CloudWatch

An IT team running a few hundred endpoints and a handful of servers often gets far with an RMM plus a network monitor. A team supporting a customer-facing application needs APM on top. MSPs care about one more thing: multi-tenancy, so each client's data stays separate and one console covers them all.

If you're comparing specific products, our roundup of IT operations tools covers 16 of them.

For monitoring-first options, the Nagios alternatives guide compares seven side by side. For the network layer on its own, see our comparison of network management software.

OpenFrame, Flamingo's open, AI-native infrastructure layer for IT and security, can run a check as a script across a client's devices and collect the output in one place, which covers the ad-hoc questions a dashboard didn't anticipate.

A Selection Checklist

Before a demo, write down what you need to see. Then make each vendor show it, on your data if they'll allow a trial.

  • Coverage. Does it reach every layer you listed, or will you need a second tool for the network or the cloud?
  • Collection. Agent for your OS mix, SNMPv3 support, WMI or WinRM, the cloud APIs you use.
  • Alerting. Duration windows, dependencies (don't alert on devices behind a dead switch), deduplication, auto-close, on-call routing.
  • Baselines. Historical retention long enough to see a full month, and trend or forecast rules for capacity.
  • Heartbeats. Can a script or job check in, and can the tool alert on a missing check-in?
  • Integrations. Tickets into your PSA or helpdesk, messages into your chat tool, an API for everything else.
  • Multi-tenancy. If you serve several clients or business units, separate views and permissions per tenant.
  • Cost model. Per device, per sensor, per metric or per host. Model it at your size now and at twice your size, because monitoring bills grow with the estate.
  • Security. How credentials are stored, what the agent can do on a device, and how the vendor handles its own breaches.

Rolling It Out Without Drowning

A new monitoring tool fails the same way every time: someone turns on every default check, the alert queue explodes, and the team learns to ignore it in a week. Roll it out in stages instead.

Start with inventory. You can't monitor what you haven't listed, so the first week is discovery: every server, network device, cloud subscription and critical service, with an owner for each.

Next, the critical path. Pick the ten things whose failure users notice first, usually internet connectivity, identity, email, file access, backups and the main business app. Set symptom checks and pages for those only.

Then capacity and hygiene. Add disk growth, certificate expiry, patch status and agent health, all routed to tickets, not pages. Let the baselines settle.

Finally, tune. After a month, review every alert that fired. Delete the ones nobody acted on, lengthen windows that flapped, and add the checks your incidents showed were missing. Put that review on the calendar every quarter, because the environment keeps changing even when the monitoring doesn't.

The Short Version

IT infrastructure monitoring works when it watches every layer, collects data the way each device allows, and only interrupts a person for something they can fix. Start from symptoms users feel, give every metric a baseline and a duration, and watch for the jobs that go quiet as closely as the servers that go down. The tool matters less than the discipline around it.

If alert noise is already the problem, start with our guide to cutting RMM and SIEM noise before adding anything new.

Conrad Lunderstedt

Conrad Lunderstedt

Solution Architect

I'm Conrad, Solution Architect at Flamingo. I've spent about 26 years in IT, roughly half of it inside MSPs and the rest in enterprise environments, so I've watched vendor decisions get made on both sides of that line. Now I spend my days talking with MSPs about the stack they already run, and helping them work through the requests and issues that come with it.

Related Content

Blog Posts

Product Releases

Podcasts

Webinars

Case Studies

Events

Onboarding Guides

Frequently Asked Questions

IT Infrastructure Monitoring

IT infrastructure monitoring is the continuous collection of health and performance data from servers, network devices, storage, cloud services, endpoints and the services running on them, with alerts when something needs a person. It covers availability, capacity and performance across every layer an application depends on.
An agent is software installed on the device that reads data locally and pushes it to a central server, which gives the deepest data and works off the office network. Agentless monitoring asks the device over the network with SNMP, WMI or WinRM, APIs or syslog. Use agents where you can install them and agentless methods for network gear, appliances and cloud services.
Use the USE method: utilization, saturation and errors for CPU, memory, disk and network. Saturation, such as a long CPU run queue or disk queue, often shows trouble before utilization maxes out. For disks, alert on the projected time until the volume is full rather than a fixed percentage.
Page only on symptoms someone can act on, such as a service being down or slow. Send trends like disk growth or expiring certificates to tickets, and keep informational events on dashboards. Add duration windows, deduplicate alerts from devices behind a failed switch, auto-close tickets when conditions clear, and review every page after the fact.
Monitoring checks known conditions on a schedule, such as whether a server is up or a disk is filling. Observability collects metrics, logs and traces detailed enough to investigate problems you did not anticipate. Infrastructure teams rely mainly on monitoring and add observability tools for complex custom applications.
Teams usually combine categories: an RMM for managed endpoints and servers, a network monitoring system for SNMP devices, an infrastructure monitor such as Zabbix or Checkmk for hosts and services, APM for applications, and cloud-native tools like Azure Monitor or CloudWatch for a single provider. Pick by the layers you need to cover.

About OpenFrame

OpenFrame isn't built to plug into your stack. It replaces it. Instead of duct-taping a dozen tools together (RMM, MDM, SIEM, patching, remote access, each its own login and bill), we bundle it into one unified platform: RMM, MDM, monitoring, automation, remote access, patch management, security monitoring, and ticketing, plus built-in AI copilots. So "does it integrate with X?" usually means: you won't need X anymore.
Most platforms give you one piece and expect you to bolt the rest on. OpenFrame unifies the whole stack in one place, with AI copilots built in. Fewer logins, fewer bills, less duct tape.

MSP AI Agents

On a five-person desk, reported deployments show $78,000 to $130,000 in annual direct labor savings, roughly 30% fewer escalations, and 15% to 20% better SLA compliance. Broader MSP adoption data adds ticket handling time cut by 45% and five to 12 points of margin, all from reclaimed capacity rather than headcount cuts.
Yes. In production MSP shops today, 10% to 25% of tickets close before a human opens them. Thread alone has processed 173 million tickets across 750-plus MSP partners at 96% triage accuracy, handing back 490,000-plus technician hours. Agents own the low-risk, high-volume work (password resets, MFA enrollment, known installs, onboarding and offboarding) and flag anything that touches production data or needs judgment for a human to take.