A security consultant spends two days on site, hands over a 40-page PDF, and three weeks later nobody can say whether the environment is better or worse than it was. The report answered "what did we find" when the question the board asked was "how exposed are we, and is that number going up or down". A security posture assessment is built to answer the second question, and this guide covers what it measures, how to score it, and how a small IT team runs one in 30 days without hiring anyone.
TL;DR
- Posture is your measured state: which controls are in place, how far they reach, and what is exposed. An audit checks evidence at a point in time. You need both, but they answer different questions.
- Score it on a scale someone else defined. NIST CSF 2.0 tiers, CIS Implementation Groups and Microsoft Secure Score each work, and each has a blind spot the post spells out.
- A score counts controls, not breach odds. Microsoft says this about its own number. Treat 88% as "88% of the recommended switches are on", nothing more.
- The 30-day run: scope and inventory, then evidence, then score and gap, then an owner and a date for every gap. Re-score quarterly and re-audit yearly.
- The first fixes are the same in almost every assessment: MFA gaps, local admin rights, unpatched internet-facing systems, backups nobody has restored.
What a Security Posture Assessment Measures
NIST's Cybersecurity Framework 2.0, published in February 2024, describes an Organizational Profile as a document that "describes an organization's current and/or target cybersecurity posture in terms of the Core's outcomes". That is the cleanest definition on offer. Posture is where you stand against a list of outcomes you chose in advance, and an assessment is the act of measuring it.
Four inputs make up the measurement. Controls coverage: which safeguards exist and how much of the estate they reach, because MFA on 50 of 100 accounts is a different number from MFA on all of them. Exposure: what an attacker can touch from outside, which is the attack surface rather than the policy. Configuration drift: settings that were right at rollout and moved since. Readiness: whether the people and playbooks would work at 4pm on a Friday.
The word to separate this from is "audit". An audit is an evidence exercise at a point in time. Someone checks that a control existed on a date, collects the proof, and files it. The security audit procedures post walks through that in full. A posture assessment is the continuous version: the same controls, measured as a score, tracked over time, and pointed at exposure rather than paperwork. The audit tells an insurer or a regulator that you did the work. The posture score tells you whether the work is holding.
Two things it is not. It is not a penetration test, which tries to break in through the gaps a posture score only counts. And it is not a compliance mapping, though the two share most of their evidence. If you need the framework side, the IT compliance guide covers who owns which evidence.
The Six Areas the Assessment Covers
Every framework carves the estate a little differently, but the same six areas show up in all of them. The table is the working checklist: what to check, what counts as evidence, and the condition that earns a pass. The evidence column matters more than it looks. A pass with no evidence behind it is an opinion, and opinions drift.
| Area | What to check | Evidence | Pass condition |
|---|---|---|---|
| Assets and inventory | Every device, account and SaaS app that touches company data | Exported device list, identity directory export, SaaS discovery | Inventory reconciles with billing and DHCP within 5% |
| Identity and access | MFA coverage, admin roles, stale accounts, shared logins | Sign-in logs, role assignments, last-logon report | MFA on 100% of accounts that can sign in, no dormant admins |
| Endpoints and patching | OS build, patch age, local admin, disk encryption, EDR agent present | RMM or Intune report, script output per device | No internet-reachable system past 30 days of patch age |
| Network and perimeter | Open inbound ports, remote access paths, VPN config, guest and IoT segmentation | External port scan, firewall rule export | No RDP, SMB or database ports open to the internet |
| Data and backup | Where data lives, who can reach it, backup coverage and last restore test | Backup job history, restore test log, sharing reports | A restore tested in the last 90 days for every critical system |
| Detection and response | Log retention, alert routing, an incident plan with names in it, last tabletop | Log config, on-call rota, plan document, exercise notes | Someone is paged for a critical alert inside 15 minutes, any hour |
Governance sits over all six. CSF 2.0 added a Govern function to hold strategy, roles and oversight, and the IT governance framework post is the long version. In an assessment it shows up as one question: does anyone own each row above, by name?
Scope the assessment before you start on the table. NIST's own five-step profile process starts with scoping, then gathering information, then creating the profile, then gap analysis, then an action plan. A first assessment can cover one office, one client, or just identity and endpoints. A narrow, finished assessment beats a wide one that stalls at row three.
How to Score It: Three Scales That Work
A score is only useful if someone else defined the scale. Three are in common use, and they measure different things.
NIST CSF 2.0 tiers. CSF says its tiers "characterize the rigor of an organization's cybersecurity risk governance and management practices". There are four: Partial (Tier 1), Risk Informed (Tier 2), Repeatable (Tier 3) and Adaptive (Tier 4). They run from ad hoc, undocumented practice to processes that are standardized, measured and improved on purpose. A tier is a judgement about how you manage risk, not a count of controls, which is why it is the scale insurers and boards recognise and the one hardest to self-assess without flattering yourself.
CIS Implementation Groups. The CIS Controls come in three tiers of safeguards. CIS defines Implementation Group 1 as "essential cyber hygiene", a foundational set of 56 safeguards for every enterprise. IG2 adds 74 more for a total of 130, aimed at organizations with dedicated IT staff and regulatory exposure. IG3 is the complete set of 153. This scale counts controls, so the score is a plain fraction: 41 of 56 IG1 safeguards in place is 73%. The cybersecurity frameworks list covers how CIS maps to NIST and ISO if you need one to satisfy another.
Microsoft Secure Score. For a Microsoft 365 estate, Secure Score is the scale you already have. Microsoft's documentation, updated in March 2026, describes each recommended action as worth 10 points or less, most scored yes or no, some scored as a percentage of coverage. Its worked example: protect 50 of 100 users with MFA and you get 5 of 10 points. The score updates as you change configuration and syncs daily. It covers Entra ID, Exchange, SharePoint, Teams, Defender and a list of third-party apps, and nothing outside them.
The r/sysadmin thread below is a good read on what the Microsoft number is and is not. The poster took an estate from 44% to 88.25% over a few months and asked whether the score is worth the pixels. The replies split between people whose insurers now ask for the number on the application form and people who have filled in those forms for years and never seen the question.
Where Scores Mislead
Every scale above has a way of reading better than the estate it describes. Five show up over and over.
License-gated points. Some of the actions that would raise a Secure Score require a license tier you do not have. The top reply in that Reddit thread makes exactly this complaint. Microsoft's docs are open about it: the full set of recommendations shows regardless of your plan, so part of your "missing" score is a price list. Score what you can act on, and log the rest as a licensing decision, not a security gap.
Scope blind spots. The same poster noted the score only covers Microsoft products, then said that was about 90% of their stack. The other 10% is where the firewall, the backup appliance and the line-of-business server live. A posture score that stops at the tenant boundary is a tenant score.
Self-assessed maturity. Tiers are a judgement, and the person judging is usually the person who built the thing. Repeatable feels true when you have a runbook. It is only true if two different people can follow it and get the same result, which is a test, not a feeling.
Counting controls instead of exposure. A control-count scale gives one point for closing an open RDP port and one point for a password policy. The open port was the whole risk. Verizon's 2025 Data Breach Investigations Report, published in April 2025, found exploitation of vulnerabilities as an initial access vector rose 34% year over year. The scale does not know which of your missing points is the one that matters. You have to.
Point-in-time drift. A score from March is a fact about March. Devices get rebuilt, a contractor gets admin for a week and keeps it, a firewall rule gets added at 11pm to make a demo work. Without a re-score, the number on the slide slowly stops describing anything.
None of this is an argument against scoring. It is an argument for reading the score with the scale's blind spot next to it, and for pairing a control count with an outside-in view of what is exposed.
How to Run One in 30 Days
This is the run sheet for a team of one to five people. It follows NIST's five profile steps, compressed to four weeks and with the evidence-gathering done by script rather than by walking around.
Week 1: scope and inventory. Decide what is in. One company, one site, or one client if you are an MSP. Then build the inventory from three sources and reconcile them: the identity directory, the device management tool, and the finance system's list of what is being paid for. The gaps between those three lists are the first findings, and they usually arrive before you have checked a single control. Pick the scale now too. CIS IG1 is the right first scale for a team that has never scored itself, because 56 safeguards is a list you can finish.
Week 2: gather evidence. Every row of the six-area table needs an export or a script output, not an interview answer. Identity: an export of accounts, MFA methods and role assignments, plus last sign-in dates. Endpoints: one script that reports OS build, last patch date, BitLocker state, local administrators group membership and whether the security agent is running, run across every device and collected centrally. OpenFrame can run that script across a client's devices and collect each device's output in one pass. Network: an external scan of your public addresses and an export of the firewall rules. Backup: the job history and the date of the last tested restore. Detection: log retention settings and who is on the pager.
Week 3: score and gap. Fill the scale. For CIS, mark each safeguard in place, partial or absent, with the evidence file named next to it. For CSF, write the current profile and pick a target profile for 12 months out. Then do the gap analysis NIST asks for: the difference between current and target, ranked. Two ranks, not one. Rank by exposure first, using the network scan and the identity export, and by effort second.
Week 4: owners and dates. Every gap gets a name and a date, in a document the business can see. Some fixes are an afternoon. Some are a quarter. Some are a licensing decision that goes to whoever holds the budget, with the points it would buy written next to the price. Set the re-score date before you close the assessment, and put a tabletop exercise on the calendar to test the readiness row, because that is the one row a script cannot measure.
If you would rather watch the shape of one before running it, Microsoft's SC-100 series has an episode on posture assessments as an architect thinks about them.
From Score to Fix: Prioritizing the Gaps
A finished assessment produces more gaps than a small team can close in a quarter. The r/sysadmin thread below is what that looks like from the inside. A posture review found that roughly 140 of 250 users had local admin, it went into the board report, and the admin got a 90-day window and no extra headcount to fix it without breaking the developers, the finance software or the legacy apps that need admin to run.
That thread is the whole prioritisation problem in one post. The finding is correct, the fix is known, and the cost of the fix lands on the people who did not create the problem. Sort gaps on two axes and the order mostly writes itself.
Exposure is the first axis: can this gap be reached from the internet, or by anyone with a stolen password? Open remote access, missing MFA, unpatched edge devices and admin rights on daily-driver accounts all score high. Effort is the second: hours, days, or a project. The high-exposure, low-effort corner goes first and is usually done inside the 30 days. MFA enforcement, closing an inbound port, disabling dormant accounts, turning on disk encryption.
High exposure and high effort is the local admin problem. It needs a plan, a tool that grants elevation per application, and a catalog of what people run. The endpoint privilege management guide covers that path. It gets a project owner and a quarter, not a weekend.
Low exposure and low effort is housekeeping. Do it in the gaps between the other work. Low exposure and high effort goes on the roadmap with a date, and the assessment says so in writing, which is what stops it becoming a surprise in next year's audit.
One more rule from that thread. Every fix that touches what users can do needs a test group and a rollback, because the security team's 90-day clock does not care that the finance software silently breaks without admin. Finding that out in a pilot of ten people costs an afternoon. Finding it out at month end costs the mandate.
Maturity Models: When Tiers Matter
A cybersecurity maturity assessment is the same exercise pointed at process rather than controls. It asks how you manage security, not just what is switched on. The CSF tiers are the common scale, and CISA's Cross-Sector Cybersecurity Performance Goals 2.0 give a checklist mapped to the CSF functions, with a self-assessment module in CISA's evaluation tool.
The tiers read like this for a 50-person company. Tier 1, Partial: patches happen when someone remembers, there is no inventory, and the incident plan is the IT lead's phone number. Tier 2, Risk Informed: leadership knows the risks and has approved priorities, but practice varies by team and by month. Tier 3, Repeatable: policies exist, are followed, are reviewed on a schedule, and two different people would run the same runbook the same way. Tier 4, Adaptive: the organization changes its practices based on what it learns from incidents, metrics and the threat landscape, and it can show the changes.
Three situations make the tier the number that matters. Cyber insurance applications increasingly ask process questions, and the cyber insurance requirements post lists the ones that move premiums. Client questionnaires, where a customer's vendor risk management process wants to know your tier before they sign. And defense supply chains, where CMMC turns maturity into a contractual requirement.
Verizon's 2025 report found that breaches involving a third party doubled to 30% of the total. That figure is why the questionnaires are getting longer, and why an MSP's own posture is now part of every client's posture.
Josh Sokol's BSides talk on measuring maturity with the CSF is the best 40 minutes on how to score tiers without fooling yourself, including how to run the scoring with people who did not build the controls.
Keeping the Posture Current
The assessment is a snapshot. Posture is a trend, and the trend needs three cadences.
Continuous for exposure. Patch age, open ports, new admin accounts and disk encryption state change weekly, so the script from week 2 runs on a schedule and the output lands somewhere a person looks. Vulnerability management software covers the scanning side; the point is that the exposure rows of the table never wait for the next assessment.
Quarterly for the score. Re-run the scale, compare to last quarter, and write two lines: what improved and what slipped. A score that goes up because you bought a license and a score that goes up because you closed 20 gaps are different stories, and the two lines are where you tell them apart.
Yearly for the audit. Once a year, someone who is not you checks the evidence. That is the audit, and it is where the posture work pays off, because the evidence files already exist and are already named.
Log retention underpins all three. Without log management in place, half the evidence rows above cannot be produced, and the detection row is a guess.
Short Version
A security posture assessment measures where you stand against outcomes you chose in advance: controls coverage, exposure, drift and readiness. Score it on a public scale, know the scale's blind spot, and treat the number as a count of switches rather than a probability. Run the first one in 30 days on one narrow scope, get evidence by script, rank gaps by exposure then effort, and put a name and a date on every one. Then re-score quarterly so the number keeps meaning something.
If the readiness row scored worst, start with the incident response plan for small teams. If it was the data row, the backup software guide covers restore testing for Windows fleets.
FAQ
What is the difference between a security posture assessment and a security audit?
An audit checks evidence that controls existed at a point in time, usually for an insurer, a regulator or a customer. A posture assessment measures the same controls as a score, adds exposure and configuration drift, and tracks the result over time. Audits prove you did the work. Posture scores show whether the work is holding.
How often should you run a security posture assessment?
Exposure checks such as patch age, open ports and new admin accounts should run continuously by script. Re-score the full scale quarterly and compare it to the previous quarter. Run a full evidence audit once a year with someone outside the team checking the files.
Is Microsoft Secure Score a good measure of security posture?
It is a good measure of how many of Microsoft's recommended actions are configured in a Microsoft 365 tenant, and Microsoft's own documentation says it is not a measure of breach likelihood. It does not see firewalls, backup appliances or non-Microsoft servers, and some points need a license you may not have. Use it as one input, next to an external scan and a control-count scale such as CIS IG1.
What is a cybersecurity maturity assessment?
It measures how an organization manages security rather than which controls are switched on. The usual scale is the NIST CSF 2.0 tiers, from Partial through Risk Informed and Repeatable to Adaptive, which NIST describes as characterizing the rigor of risk governance and management practices. Insurers, customer questionnaires and defense contracts are the three places it tends to be asked for.
Which framework should a small team use for a first assessment?
CIS Implementation Group 1. It is 56 safeguards that CIS defines as essential cyber hygiene, it is finite, and the score is a plain fraction. Move to a CSF 2.0 profile once IG1 is mostly in place and someone outside IT wants a maturity answer.
What should be fixed first after a posture assessment?
Anything reachable from the internet or by a stolen password that can be closed in hours: MFA gaps, open remote access ports, dormant accounts, missing disk encryption. Then the high-exposure projects, such as removing local admin rights, with an owner, a pilot group and a quarter to do it.

Aliaska Varieva
Head of Platform
Hi! I’m Aliaska, and I’ve been working as a software engineer (mostly Java + a bit Kotlin) for over 8 years now. I mostly spend my time building backend services, integrating systems, fixing bugs (the fun part 🙃), and making sure things don’t fall apart behind the scenes.
