Updated: October 2026
The alert fires at 4:52 on a Friday, and the first question in the room is who's in charge. A small IT team doesn't need a 60-page binder to answer that, but it does need roles, a first-hour routine and a few playbooks decided in advance. This is a working guide to incident response for teams of two to ten people, built on NIST's 2025 guidance and three incidents small teams meet again and again.
TL;DR
- Incident response is how you detect, contain, fix and learn from a security incident. NIST's April 2025 revision of SP 800-61 ties it to the six CSF 2.0 functions instead of a separate four-phase circle.
- On a small team, name four roles before anything happens: incident lead, technical lead, communications and scribe. Keep the outside contacts (insurer, legal, MSP, bank, law enforcement) on one page you can reach offline.
- The first hour decides the cost. Confirm, contain, preserve evidence, call the insurer, and move the team to a channel the attacker can't read.
- Write three playbooks first: ransomware, business email compromise and a lost laptop.
- Notification clocks start early: 72 hours under GDPR, four business days after a materiality call for SEC registrants, and whatever your cyber insurance policy says.
What Incident Response Means
Plenty happens on a network that isn't an incident. A failed login is an event. A thousand failed logins against one account, followed by a success from another country, is an incident.
NIST draws the same line. An event is any observable occurrence. An incident is an occurrence that "actually or imminently jeopardizes" the confidentiality, integrity or availability of information or systems, or breaks security policy.
Incident response is everything the team does from the moment that line is crossed. It covers detection, triage, containment, removing the attacker, recovery, communication and the review afterwards. It overlaps with two neighbors: the service desk's incident management process, which handles outages and broken printers, and disaster recovery, which rebuilds systems after the damage.
If you're sorting out which plan covers what, our guide on BC vs DR vs IR plans draws the borders.
This post stays on the security side: someone got in, or tried to, and you have to deal with it.
The Incident Response Lifecycle, Then and Now
For more than a decade, the reference model came from NIST SP 800-61 Rev. 2, published in 2012. It's a circle with four phases:
| Phase | What happens |
|---|---|
| Preparation | Tools, training, contacts, playbooks |
| Detection and Analysis | Spot it, confirm it, scope it |
| Containment, Eradication and Recovery | Stop the spread, remove the attacker, restore service |
| Post-Incident Activity | Lessons learned, fixes, report |
In April 2025, NIST replaced it. SP 800-61 Rev. 3 is written as a CSF 2.0 Community Profile, and it maps incident response onto the six CSF functions: Govern, Identify, Protect, Detect, Respond and Recover.
The reason is practical. NIST notes that in 2012 incidents were rarer and recovery usually took a day or two. Today recovery often takes weeks or months, so a separate team running a separate circle no longer fits. In the new model, Govern, Identify and Protect are the preparation that happens all year. Detect, Respond and Recover are the incident response itself. Lessons learned feed an Improvement category that informs every function, and NIST says they should be shared as soon as they're found, not after recovery ends.
For a small team, the four phases still work fine as a checklist. The useful change from Rev. 3 is the mindset: preparation is part of normal operations, and you improve during the incident, not only in a meeting three weeks later. NIST itself says to use whichever life cycle model suits you.
Who Does What on a Small Team
On a team of three, one person may hold two roles. That's fine. What breaks an incident is nobody holding a role, or two people both acting as the lead.
| Role | Owns | On a 3-person team |
|---|---|---|
| Incident lead | Decisions, severity, priorities, when to escalate or call it closed | IT manager or senior tech |
| Technical lead | Investigation, containment, eradication, recovery work | The most hands-on tech |
| Communications | Updates to leadership, staff, customers and outside parties | IT manager, or a business owner |
| Scribe | Timeline, decisions, evidence log | Whoever is left, even part time |
The scribe sounds optional. It isn't. The timeline becomes your insurance claim, your regulator notice and your review document. Write down what happened, when, who decided what, and why.
Then there's the outside world. NIST's revision points out that incident response now involves many third parties, including managed security providers doing the primary work. Put these contacts on one page:
| Contact | Why you call them | When |
|---|---|---|
| Cyber insurer's breach hotline | Coverage often depends on calling first and using their panel firms | First hour, before hiring anyone |
| Legal counsel | Privilege, notification duties, law enforcement contact | First hours |
| MSP or MSSP | Hands, tools and logs you don't have | As soon as severity is set |
| Bank's fraud team | Recalling a fraudulent transfer | Immediately on any payment fraud |
| Law enforcement (FBI IC3, local field office) | Fund freezes, intelligence, sometimes decryptors | Within hours for fraud and ransomware |
| Key vendors | Your email host, identity provider, line-of-business apps | When their systems are in scope |
Keep this page somewhere the incident can't reach. If the plan lives only on the file server and the file server is encrypted, you don't have a plan.
How Incidents Get Noticed
Small teams rarely spot incidents from a dashboard at 3 a.m. They hear about them. The signal arrives from five places, and the plan should say who watches each one.
Users are the first source. "My mailbox sent emails I didn't write" or "all my files have a strange extension" is a detection, and staff should know the one number or address to report it. Tools come second: endpoint alerts, risky sign-in flags in Entra ID, backup jobs that suddenly fail, firewall alerts. Your MSP or security provider is third, if you have one. Fourth are outsiders: a supplier who got a suspicious invoice from your domain, a bank that flags a payment, or law enforcement. Last is the slow signal nobody likes: the disk that filled overnight, the server that's suddenly busy at 2 a.m.
Whatever the source, triage asks the same three questions. Is it real? What's affected: one account, one machine, or something shared? Is it still happening? Those answers set the severity, and the severity sets everything else.
Severity Levels That Drive the Response
Severity decides how loud the response gets: who gets woken up, whether you call the insurer and how often you update leadership. Keep it to four levels and define them by impact, not by how scary the alert name sounds.
| Severity | Looks like | Response |
|---|---|---|
| SEV 1 | Active ransomware, confirmed data theft, attacker with admin rights | All hands, insurer and legal now, leadership updates hourly |
| SEV 2 | One compromised account with sensitive access, malware on a server | Incident lead plus technical lead, same day, leadership informed |
| SEV 3 | Malware caught and quarantined on one PC, phishing clicked without credential entry | Technical lead during business hours, logged |
| SEV 4 | Suspicious event that turned out benign | Logged and closed |
Two rules keep this working in practice. Anyone can raise the severity, but only the incident lead lowers it. And when you're unsure between two levels, pick the higher one and step down once the facts come in.
If your service desk already uses a priority matrix for outages, reuse the shape. The same impact-and-urgency thinking applies; security incidents just get their own labels so nobody confuses a SEV 1 breach with a P1 printer outage.
The First Hour Checklist
The first hour is where small teams win or lose. Here's the order that holds up across incident types:
- Confirm it's real. Check the alert against a second source: a sign-in log, the endpoint tool, the user.
- Name the incident lead and set a severity. Say it out loud or type it in the channel.
- Move to an out-of-band channel. If email or Teams might be compromised, use phones, Signal or a separate tenant.
- Contain the spread, not the evidence. Isolate machines from the network instead of switching them off. Disable accounts instead of deleting them.
- Preserve evidence. Export logs before they roll over, and don't reimage anything yet.
- Call the insurer's hotline. Many policies require notice before you hire a forensics firm or negotiate with anyone.
- Start the timeline. The scribe logs every action with a timestamp from here on.
- Brief leadership. What happened, what's affected, what you're doing, when the next update comes.
Isolation beats shutdown because memory holds evidence: running processes, network connections, sometimes the encryption keys. Pulling the network cable, or isolating the device from your endpoint console, stops the spread and keeps that evidence.
Containment is a trade-off, and the incident lead owns it. Pulling a server off the network stops the spread but may stop payroll too. Disabling the CEO's account stops the attacker but also the CEO. Decide by asking what's worse in the next hour: more damage, or more downtime. Write the answer and the reason into the timeline, because someone will ask later.
Three Playbooks to Write First
A playbook is the incident-specific version of the checklist: who does what, in which order, with which commands. Write these three before anything fancier. Keep each to a page.
Ransomware
Ransomware is the classic SEV 1, and the one where order matters most. The attack usually starts days before the encryption, so the playbook has to look backwards as well as stop the damage. Our walkthrough of how a ransomware attack moves covers that timeline.
| Step | Action |
|---|---|
| Contain | Isolate affected devices, disable the accounts the attacker used, block the remote access path (VPN, RDP, remote tool) |
| Protect backups | Confirm backups are offline or immutable and cut backup server access from the production network |
| Scope | Find patient zero, check for data theft (large uploads, archive tools, cloud sync), list affected systems |
| Call | Insurer hotline, legal, law enforcement; don't contact the attacker without counsel |
| Eradicate | Reset all privileged credentials, including service accounts and the Kerberos KRBTGT account twice |
| Recover | Rebuild from known-good images, restore data from clean backups, verify before reconnecting |
The backup step decides the outcome. If backups were reachable from the same admin accounts the attacker stole, they may be gone too. That's why the restore itself should be rehearsed; our guide to disaster recovery testing covers the five test types.
Business Email Compromise and Account Takeover
Business email compromise is where the money goes. The FBI's IC3 recorded 24,768 BEC complaints and $3.05 billion in reported losses in 2025, according to its 2025 annual report.
Speed matters more here than anywhere else. In 2025, the IC3 Recovery Asset Team froze $679 million of $1.16 billion in attempted theft, a 58% success rate, and it can only freeze money that hasn't moved on yet. If a fraudulent payment went out, calling the bank to request a recall comes before anything technical.
Then take back the account. Microsoft's guide to responding to a compromised email account sets the order:
| Step | Action |
|---|---|
| Disable or reset | Disable the account during the investigation; if you can't, reset the password and don't send the new one by email |
| Revoke sessions | Revoke-MgUserSignInSession invalidates active sessions and refresh tokens |
| Check MFA methods | Remove authenticator devices or phone numbers the attacker added |
| Check app consents | Revoke OAuth apps the user didn't approve |
| Check roles | Remove admin roles that shouldn't be there |
| Check forwarding and rules | Get-InboxRule -IncludeHidden finds rules that move or forward mail out of sight |
After that, look backwards. Pull sign-in logs for the account, trace the messages it sent in the window, and warn the customers and suppliers who received them. Attackers who control a mailbox often send invoices with new bank details to real contacts, so the second victim may be someone else's finance team.
This r/msp thread is a good reminder that an incident can surface on day one of a new engagement, and that the playbook has to work on an environment you didn't build.
Lost or Stolen Laptop
A lost laptop often turns out to be minor, as long as two things were true before it went missing: the disk was encrypted and the device was managed.
| Step | Action |
|---|---|
| Confirm | When and where was it last seen, was it signed in, was it locked |
| Revoke | Revoke the user's sessions and reset their password; remove the device from trusted device lists |
| Wipe | Send a remote wipe or lock from your device management tool |
| Check encryption | Confirm the device reported BitLocker or FileVault as on, and that the recovery key is escrowed |
| Assess data | Work out what was stored locally and whether any of it is regulated personal data |
| Decide on notification | With counsel, check whether breach laws apply; some regimes treat encrypted data differently |
If encryption status can't be proven, treat it as a data exposure until counsel says otherwise.
Evidence and Logging
Evidence does three jobs: it tells you what happened, it supports the insurance claim, and it backs up any notification you make. Small teams lose it in two ways. They reimage before anyone looked, or the logs they need rolled over before anyone exported them.
Check your retention before you need it. In Microsoft 365, Audit (Standard) keeps records for 180 days by default for logs generated since 17 October 2023, according to Microsoft Purview's documentation. Audit (Premium) keeps Exchange, SharePoint, OneDrive and Entra records for a year. Firewalls, VPN concentrators and endpoint tools often keep far less, sometimes days.
Preserve in this order: volatile data first (memory, running processes, network connections), then logs from the systems involved, then disk images if forensics will need them. Keep a simple chain-of-custody note for each item: what it is, who collected it, when, and where it's stored. Make sure your devices agree on the time, because a timeline built from logs with drifting clocks lies to you.
A central log store helps more than any single tool. If sign-ins, endpoint alerts and firewall logs land in one place, the technical lead can build the timeline in an hour instead of a day. That's the job a SIEM does for larger teams. In OpenFrame, you can run the collection commands as a script across a client's devices and gather the output in one place, which saves the log-by-log trips when several machines are in scope.
Communication and Notification
Communication runs on two tracks. Inside, leadership needs short updates at a set rhythm: what we know, what we don't, what we're doing, next update at a set time. Staff need to know what to do and what not to do, usually "don't use email for anything sensitive" and "don't talk about this outside." Say less, more often, and never guess out loud.
Outside, several clocks may already be running. Name them only after counsel confirms which apply:
- GDPR. Under Article 33, a controller notifies the supervisory authority "not later than 72 hours" after becoming aware of a personal data breach, unless the risk to people is unlikely. A processor must tell its controller "without undue delay."
- SEC registrants. Public companies disclose a material cybersecurity incident on Form 8-K Item 1.05, generally within four business days after determining it's material.
- Contracts and insurance. Customer contracts, data processing agreements and your cyber policy often set their own notice windows. They are the easiest deadlines to miss because nobody reads them until the day they matter.
- US state breach laws. Every state has one, with different triggers and timelines. This is counsel's call, not IT's.
The practical rule for IT: log the moment you became aware, because many clocks start there, and hand the notification decision to legal with the facts ready.
Recovery and the Post-Incident Review
Recovery ends when systems are clean and trusted again, not when they're back online. Rebuild compromised machines rather than cleaning them, restore data from backups taken before the attacker's first known activity, and watch the environment closely for a few weeks. Attackers who kept a second way in tend to come back.
Then hold the review within two weeks, while people remember. Keep it blameless: the goal is to fix the system, not to find a person. Walk the timeline, and answer four questions. What happened? What went well? What slowed us down? What changes, who owns each change, and by when?
The review is also where the report comes from. One incident responder who says they've handled around a thousand incidents shared the report template they use to get clients to act on findings:
NIST's 2025 revision adds one habit worth copying: don't wait for the review to share a lesson. If the technical lead finds that the backup server was reachable with a normal admin account, fix it that day.
When to Bring In Outside Help
Some incidents are bigger than a small team, and knowing that early saves money. Call a digital forensics and incident response (DFIR) firm, usually through your insurer's panel, when any of these is true: a domain admin or global admin account was compromised, you suspect data was taken, regulated data (health, payment, personal data under GDPR) is involved, or the attacker is still active and you can't see how they got in.
A retainer agreed in advance beats a cold call on a Friday night. It fixes the rates, the paperwork and the response time before you need them, and some insurers will name the firms they already accept.
Practice It Before You Need It
A plan nobody has rehearsed is a guess. The cheapest rehearsal is a tabletop exercise: an hour around a table, one scenario, and a facilitator who adds a new fact every ten minutes. "The insurer's hotline goes to voicemail." "The CFO has already replied to the attacker." The gaps show up fast.
Run one a year at minimum, rotate the scenario between the three playbooks, and feed the findings into the plan the same week.
Professor Messer's Security+ lesson on incident response is a short refresher on the classic phases for anyone new to the team:
A One-Page Incident Response Plan Template
This is the page to print, pin in the team channel and store offline. Fill in the right column and review it every six months.
| Section | What to write |
|---|---|
| Scope | Which systems, sites and data this plan covers |
| Roles | Incident lead, technical lead, communications, scribe, each with a named backup |
| Severity | Your four levels with one example each |
| First hour | The eight-step checklist above, adjusted to your tools |
| Contacts | Insurer hotline and policy number, legal, MSP, bank fraud line, IC3 and local FBI field office, key vendors |
| Out-of-band channel | Where the team meets if email and chat are compromised |
| Evidence | Where logs live, how long they're kept, who exports them |
| Playbooks | Links or printed pages for ransomware, BEC and lost laptop |
| Notification | Who decides, which regimes likely apply, where the awareness time is logged |
| Review | Who runs the post-incident review and where findings are tracked |
IBM Technology's short explainer on the response side of security architecture is a good companion for the leadership conversation about why this page exists:
Incident Response, in Short
Incident response for a small team comes down to decisions made before the alert: who leads, what counts as severe, what happens in the first hour, and who gets called. Write the three playbooks, keep the contacts offline, check how long your logs live, and rehearse once a year. NIST's 2025 guidance makes the same point in bigger words: preparation is ongoing work, and every incident should make the next one smaller.
Next, see how the same thinking applies after the damage, in our guide to disaster recovery testing.

Aliaska Varieva
Head of Platform
Hi! I’m Aliaska, and I’ve been working as a software engineer (mostly Java + a bit Kotlin) for over 8 years now. I mostly spend my time building backend services, integrating systems, fixing bugs (the fun part 🙃), and making sure things don’t fall apart behind the scenes.
