Security Maturity Assessment: When and How to Perform One

The US Department of Energy's Cybersecurity Capability Maturity Model contains 356 cybersecurity practices grouped into 10 domains, and the department states that this model and its tool "were designed to enable any organization to complete a self-evaluation in a single day." Both halves of that sentence are worth sitting with. A defensible view of a security program is measured in a few hundred specific practices rather than a general impression, and at least one government model is built so that the first pass through them does not need a quarter.
A security maturity assessment is the exercise that turns that raw material into a current, scored view of a security program. It is not a vulnerability scan and not a compliance audit. It measures how developed, repeatable and consistently governed your security practices are, domain by domain, against a published model, and it produces a baseline you can re-measure against later.
The question that decides whether the exercise is worth doing is not "how" but "when." An assessment run because a board slide needs a chart produces a chart. An assessment run because a specific decision is waiting on it produces a roadmap someone will fund. This guide covers both halves: the triggers that make a security maturity assessment the right next move, the models you can score against and what each one is actually built for, the step-by-step process, and what to do with the levels once you have them.
Key takeaways
- A maturity assessment measures capability, not compliance. An audit asks whether a control existed and operated over a defined period; a maturity assessment asks how consistently, how repeatably and how well governed that control is.
- The trigger determines the scope. Investor diligence, an acquisition, a new framework entering scope, a failed customer questionnaire and a post-incident review each demand a different boundary, a different model and a different depth of evidence.
- NIST CSF 2.0 Tiers are not a five-level maturity ladder, and treating them as one produces a score the framework does not support. The Tiers characterize the rigor of risk governance and risk management across a Profile, not the maturity of each individual control.
- Uniform high maturity is the wrong target. The C2M2 states plainly that striving for the highest level in every domain may not be optimal, and that costs should be weighed against benefits per domain.
- A score is only as good as the artifact behind it. Four artifacts separate the levels: the document that defines a practice, the record that it ran, the measurement of how well, and the evidence that the measurement changed something.
What is a security maturity assessment?
A security maturity assessment is a structured evaluation of how developed and consistently applied an organization's security practices are, scored against the levels defined in a published maturity model and reported by domain rather than as a single number. It looks at governance, process, documentation, ownership and measurement alongside the technical controls themselves, and its output is a current-state baseline plus the gap to a defined target state.
What the exercise is called signals how far it reaches rather than what it does. Read a "cybersecurity" framing as stopping at the technology estate, and an information security maturity assessment as reaching further, into physical security, personnel security and business continuity. Settle which of those you mean in the scoping conversation, because the label alone will not settle it. What does not change with the label is the measurement question. Not "do you have a control?" but "how reliably does that control behave when nobody is watching it?"
That question is what separates maturity from posture. Posture is a point-in-time state: which ports are open, which endpoints are unpatched, which accounts are over-privileged. Maturity is a property of the process that produced that state. Two organizations can present identical scan results while one of them got there through a documented, owned, measured patch process and the other got there because a capable engineer spent the weekend on it. The first result is repeatable. The second is a person.
Maturity models exist because that distinction needs a vocabulary. Each model defines a small number of named levels, describes what practice at each level looks like, and gives assessors a shared way to say "this is documented but not measured" without inventing a scale from scratch.
Cyber maturity assessment vs. risk assessment, audit, and penetration test
These four exercises are easy to commission interchangeably, and they answer entirely different questions. Each produces a report that looks authoritative, and none of them substitutes for the others, so the choice made at the start determines whether the output can answer the question that prompted it.
| Exercise | The question it answers | Output | What it cannot tell you |
|---|---|---|---|
| Security maturity assessment | How developed and repeatable are our security practices? | Scored levels per domain, current vs. target, prioritized gaps | Whether a specific system is exploitable today |
| Risk assessment | Which threats to which assets carry what likelihood and impact? | Ranked risk register with treatment decisions | How well the controls treating those risks are run |
| Audit or attestation | Did defined controls exist and operate over a defined period? | Opinion or report against stated criteria, with exceptions | Anything about controls outside the audit scope |
| Penetration test | Can an attacker achieve a defined objective in this environment? | Findings with proof of exploitation and remediation advice | Whether the finding is a one-off or a process failure |
The relationship worth understanding is the one between the maturity assessment and the audit. A SOC 2 engagement is an attestation performed by a CPA firm against stated criteria over a stated audit period, and ISO/IEC 27001 is a certification issued by an accredited certification body. Both are binary at the boundary: you are in scope or out, the control operated or it did not. A maturity assessment has no pass mark. It is the instrument you use before the audit to find out what the audit will say, and after the audit to decide which of the exceptions are symptoms of the same missing process.
The mechanics of scoping, evidence collection and reporting are covered step by step in our guide to running an internal security audit.
When to perform a security maturity assessment
Of the eighteen pages currently ranking for this topic whose text we could read in full, nine answer the question with a re-assessment cadence. Not one treats a financing event as a trigger, and not one maps a trigger to the scope and the model it calls for, which is the part that decides what you are actually commissioning. A cadence on its own is the weakest available reason to run an assessment, because one with no decision waiting on it produces a document nobody reads. The stronger reasons are events that change what someone needs to know about your security program, and each of them changes the scope, the model and the depth of evidence required.
The table below is the artifact to take into the scoping conversation. Find the row that describes why the subject came up, and the last three columns tell you what you are actually commissioning.
| Trigger | The question behind it | Scope it to | Model that fits |
|---|---|---|---|
| Funding round or investor diligence | Is security a liability that changes the price or the terms? | Whatever the data room will expose, plus the production environment | NIST CSF 2.0 Organizational Profile |
| Buy-side or sell-side M&A | What are we inheriting, and what does it cost to fix? | The target's full estate, including shadow and legacy systems | NIST CSF 2.0, or the acquirer's own model for comparability |
| New framework or regulation in scope | How far are we from the bar, and in what order do we close it? | The systems and data the framework reaches | The framework's own model: CMMC, CIS Controls IGs, or the CSF |
| Customer security questionnaires | Why do we keep losing deals at the security review? | The product, its hosting environment and the supporting processes | CIS Controls Implementation Groups |
| After an incident | Was this bad luck or a process that was always going to fail? | The domains the incident touched, then their upstream dependencies | C2M2 for the affected domains |
| AI adoption | Do our existing controls reach the systems we just deployed? | Data flows, model access, logging and third-party model use | NIST AI RMF alongside your security model |
| New security leadership or annual cadence | Where is the program, and what should the next budget buy? | The whole program, scored by domain | C2M2 or the full CSF Core |
Several of those rows deserve a note on how they differ from the obvious reading.
- Transaction triggers are about disclosure, not reassurance. In investor diligence or an acquisition, the assessment's job is to put every known weakness on the table with a cost and a date attached, before someone else finds it and prices it for you. A maturity report that presents a clean picture and is later contradicted by diligence is worse than no report. Security has moved from a footnote in the IT section to a workstream with its own findings, and the evidence a buyer reads is covered in our analysis of how SOC 2 and security posture change M&A due diligence.
- A framework entering scope is a sequencing problem, not a gap list. Anyone can produce the list of requirements you do not meet. The reason to score maturity first is that requirements cluster: a single missing process, such as nobody owning access reviews, will surface as a dozen separate findings across a dozen control families. Scoring by domain shows you the one fix behind the twelve findings.
- Questionnaire failures can be a measurement of your evidence rather than your controls. Where the same questions keep costing deals, check first whether the control exists but nothing written down proves it. That is a documentation and governance maturity gap, and it sits at the step the models describe as the move from "performed" to "managed."
- Post-incident, the value is in the domains the incident did not touch. An incident review narrows onto the failure path by design. A maturity assessment run alongside one asks the wider question: which other domains are running at the same level as the one that just failed?
On cadence: an annual full assessment with a shorter re-score of the domains under active remediation is a defensible default, but the cadence should follow the roadmap rather than the calendar. Re-measure a domain when the work you funded to change it has actually landed, not three months later because the fiscal year ended.
When AI adoption is the trigger
AI deployment reaches a security program as a product decision rather than a security decision, which is why it is easy to miss as a trigger. The relevant change is that a new class of system now reaches production data through paths your existing control set was not scoped to cover: model inputs, retrieval indexes, agent permissions and third-party inference endpoints.
Two things make this a scoping question rather than a new assessment. First, a general security maturity model will not ask about model governance, so the AI systems will score as invisible rather than as immature. Second, the AI-specific framework has its own maturity vocabulary that has to be reconciled with your security one.
We cover scoring against the AI framework in Scoring Your Organization Against NIST AI RMF, and the wider organizational readiness dimensions in The AI Readiness Checklist.
Cybersecurity maturity levels explained
Published models put the number of levels at three, four or five and give them different names, but the progression underneath is recognizably the same one. What moves a domain up is not the presence of more technology. It is the addition of documentation, then of ownership and policy, then of measurement, then of the feedback loop that acts on the measurement.
The five-level version below is a composite. It is not any single model's scale, and no assessment should be scored against it without the scoring definitions being written down first.
Read the ladder as a description of what breaks when the level is not reached.
| Level | Common labels | What practice looks like | What breaks below this |
|---|---|---|---|
| 1 | Initial, Partial, Ad hoc | The work happens when someone notices it needs doing | Nothing is predictable; outcomes track individual effort |
| 2 | Repeatable, Developing, Managed | The work is documented and resourced, and happens the same way twice | Knowledge leaves when the person leaves |
| 3 | Defined | Policy directs the work, accountability is assigned, and skills are matched to it | Practice drifts between teams with no authority to settle it |
| 4 | Quantitatively managed, Measured | Effectiveness is tracked against defined metrics, not assumed | You cannot tell a control that works from one that is merely present |
| 5 | Optimized, Adaptive | Measurement feeds change, and the program adapts to new threats and lessons learned | Improvements are reactions to incidents rather than to evidence |
Two cautions about this ladder, both of which come from the models themselves rather than from convention.
The first is that levels belong to domains, not to organizations. The C2M2 is explicit: its maturity indicator levels "apply independently to each domain in the model," so that "an organization could be operating at MIL1 in one domain, MIL2 in another domain, and MIL3 in a third domain." A single organization-wide maturity number is an average of things that should not be averaged, and it hides exactly the variation the assessment exists to expose.
The second is that the top of the ladder is not the target. The C2M2 states that "striving to achieve the highest MIL in all domains may not be optimal" and that organizations "should evaluate the costs of achieving a specific MIL versus its potential benefits." It does set one floor: the model "was designed so that all companies, regardless of size, should be able to achieve MIL1 across all domains." A defensible target is a per-domain decision that follows your risk, your obligations and your budget, with a level assigned to each domain deliberately rather than aspirationally.
Cybersecurity maturity models compared
A cybersecurity maturity model is a published scale of practice, from absent through documented to measured and self-correcting, against which an organization's own practice can be placed. Its name varies with its reach, exactly as the name of the assessment does: a security maturity model, a cyber maturity model, an information security maturity model and an IT security maturity model all describe that same instrument.
What is not interchangeable is the purpose each one was built for. The five below are the cybersecurity maturity models worth knowing, and scoring against one built for a different purpose is how assessments produce levels that nobody can act on.
| Model | Published by | Level structure | Built for |
|---|---|---|---|
| NIST CSF 2.0 | NIST | Four Tiers applied to a Profile, not a per-control ladder | Describing current and target cybersecurity outcomes in any sector |
| C2M2 v2.1 | US Department of Energy | MIL0 to MIL3, independently per domain | Self-evaluation across IT and OT, including critical infrastructure |
| CMMC | US Department of Defense | Three levels: self-assessment at Level 1, self or third-party at Level 2, government-led at Level 3 | Contractual assurance for defense contractors handling FCI and CUI |
| CIS Controls v8.1 Implementation Groups | Center for Internet Security | IG1, IG2, IG3, cumulative | Prioritizing which safeguards to implement first, by enterprise profile |
| CISA Zero Trust Maturity Model v2 | CISA | Traditional, Initial, Advanced, Optimal, per pillar | Planning a zero trust architecture across five pillars |
NIST Cybersecurity Framework maturity assessment: Profiles do the work, Tiers do not
Two of the pages we could read in full put the four CSF Tiers under the heading "NIST CSF maturity levels," and a third carries that phrase as the subject of its title. Among the eighteen readable pages, not one corrects the point. The framework does not describe the Tiers that way.
In the NIST Cybersecurity Framework (CSF) 2.0, the Tiers "characterize the rigor of an organization's cybersecurity risk governance and management practices, and they provide context for how an organization views cybersecurity risks and the processes in place to manage those risks." They run Partial, Risk Informed, Repeatable and Adaptive, and they apply to a Profile as a whole. The framework adds that Tiers "should complement an organization's cybersecurity risk management methodology rather than replace it."
The instrument CSF 2.0 gives you for the actual assessment is the Organizational Profile, and its five documented steps map almost exactly onto a maturity assessment: scope the Profile, gather the information needed, create the Profile, "analyze the gaps between the Current and Target Profiles, and create an action plan," then implement the action plan and update the Profile. A Current Profile "specifies the Core outcomes that an organization is currently achieving (or attempting to achieve) and characterizes how or to what extent each outcome is being achieved." That last clause is where maturity lives, and it is deliberately left to you to define the scale.
The practical consequence is that applying a five-level scale to CSF outcomes is legitimate, but the scale is yours and has to be defined and published with the results. Presenting it as "the NIST maturity levels" attributes a scale to NIST that NIST did not publish. A NIST Cybersecurity Framework maturity assessment is therefore two artifacts rather than one: the Current and Target Profiles the framework defines, and the scoring scale you author to sit underneath them.
C2M2: cumulative levels and a DOE self-evaluation tool
The Cybersecurity Capability Maturity Model ships with a self-evaluation tool of its own, published by the Department of Energy alongside the model document. Its 356 practices sit across 10 domains, its MIL0 to MIL3 levels are cumulative within each domain, and the department notes that the model "is designed for use in both Information Technology (IT) and Operational Technology (OT) environments and aligns with the National Institute of Standards and Technology's (NIST's) Cybersecurity Framework (CSF)." That IT and OT reach is why it is the model to reach for when an assessment has to cross into plant, field or fleet systems that a purely corporate cybersecurity maturity model was never written for.
The cumulative rule matters for scoring discipline. To claim MIL2 in a domain you must perform every practice at MIL1 and MIL2 in that domain; MIL3 requires all three levels. There is no partial credit for strong practice at a high level while a foundational practice is missing, which is precisely the failure mode that flatters organizations with good tooling and weak process.
CMMC and CIS Implementation Groups: models with an external bar
Two of these answer to somebody outside the organization. The CMMC Program, established in the Department of Defense's Cybersecurity Maturity Model Certification (CMMC) Program final rule, "provides for assessment at three levels: basic safeguarding of FCI at CMMC Level 1, broad protection of CUI at CMMC Level 2, and enhanced protection of CUI against risk from Advanced Persistent Threats (APTs) at CMMC Level 3." Those levels are contractual requirements tied to what information a contract involves, not a ladder you climb at your own pace.
Who performs the assessment varies by level, and CMMC is worth stating precisely here because a summary of it as an externally assessed model is not accurate. The rule sets out a Level 1 self-assessment and a Level 2 self-assessment conducted by the organization itself, a Level 2 certification assessment conducted by an authorized third-party assessment organization, and a Level 3 certification assessment conducted by the government. Only the last two are external.
The CIS Controls Implementation Groups work as a prioritization device rather than a maturity scale. CIS defines IG1 as "essential cyber hygiene," the foundational set of safeguards it says every enterprise should apply; IG2 builds on IG1, and IG3 comprises all the Controls and Safeguards. Scoring an organization against IG1 answers a narrower question than a five-level ladder does, and a more immediately actionable one: are the basics in place, yes or no. CIS also publishes an assessment tool of its own, CIS CSAT, which it describes as a way to "Assess and measure Controls implementation."
If you are choosing between frameworks for the first time, our guide to cybersecurity standards and frameworks sets out what each one covers and who it applies to.
Cyber resilience maturity assessment: measuring recovery, not just prevention
A cyber resilience maturity assessment narrows the same method onto the question of what happens after a control fails: detection, response, recovery and the continuity of the business processes that depend on the affected systems. It is a scope choice rather than a separate discipline, and it does not require a separate cyber resilience maturity model. In CSF 2.0 terms it concentrates on the DETECT, RESPOND and RECOVER Functions; in C2M2 terms on the domain the model names Event and Incident Response, Continuity of Operations.
The reason to run it as a distinct exercise is that prevention maturity and recovery maturity diverge. A program can hold high maturity in identity, vulnerability management and platform security while its recovery practices have never been exercised end to end. The CISA Zero Trust Maturity Model offers a useful structural precedent here: it scores maturity separately per pillar rather than as one figure, and it describes a journey advancing "from a Traditional starting point to Initial, Advanced, and Optimal," with each stage requiring "greater levels of protection, detail, and complexity for adoption."
How to perform a cybersecurity maturity assessment step by step
The sequence below assumes a trigger has been identified and a model chosen. The order matters more than the calendar: every step depends on a decision made in the one before it, and an assessment that skips the scoping conversation will run long and still answer nobody's question.
1. Write down the decision the assessment has to inform
Before scoping anything, name the decision in one sentence: the budget being set, the deal being closed, the certification being pursued, the incident being answered for. That sentence determines everything downstream, and it is the test you apply at the end to decide whether the report is finished.
2. Set the boundary, and write down what is outside it
Scope by business function and data flow rather than by org chart. Name the entities, environments, business units and third-party platforms that are inside, and list what is deliberately outside with the reason. An unstated exclusion is what lets a maturity report be contradicted later, because a reader who is not told otherwise assumes the score covers everything.
Where subsidiaries, acquired companies or OT environments are involved, decide early whether they are scored separately or rolled in. Rolling a recently acquired environment into a group score flattens exactly the signal the assessment was commissioned to find.
3. Choose the model and define the scale in writing
Pick from the table above based on the trigger, then write down the scoring scale you will use, including what evidence moves a domain from one level to the next. C2M2 and the CIS Implementation Groups come with their scale and its definitions already written. The CSF does not, for the reason set out in the previous chapter, so if you are attaching your own levels to CSF outcomes, that scale has to be authored here and published with the results.
Resist the urge to score against two models at once. Map between them afterwards if you need both views; scoring against both simultaneously doubles the work and produces two sets of numbers that disagree at the edges.
4. Set the target state before you measure the current one
Setting targets after the measurement biases the whole exercise, so do it before. Assign a target level to each domain first, with a reason tied to risk, obligation or customer requirement. If you measure first, the targets get anchored to the findings, and everything conveniently lands one level above wherever you happen to be.
Targets should differ by domain. Identity and access management in a company holding regulated data may warrant the top of the scale; physical security in a fully remote company may not warrant more than the floor.
5. Gather evidence, not opinions
This is where security maturity assessments earn or lose their credibility. A level assigned from an interview answer is an opinion about a control. A level assigned from an artifact is a measurement, and it can be re-checked by someone else next year.
For each domain, collect in roughly this order:
- The document that defines the practice. Policy, standard or runbook, with its approval date and owner.
- The record that it ran. Tickets, logs, review sign-offs, meeting minutes, exported configuration states.
- The measurement of how well it ran. Coverage percentages, time-to-close figures, exception counts, whatever the process itself produces.
- The evidence that the measurement changed something. A decision, a re-prioritization, a budget change traceable to the metric.
Those four artifacts map onto the ladder: the first supports level 2, the second level 3, the third level 4, the fourth level 5. A domain that can produce the document defining the practice but no record of it ever running is level 2, no matter how confident the interview was. Interviews are still worth doing, but their job is to find out where the artifacts live and to surface the practices nobody documented, not to set the score.
6. Score by domain, and record the evidence against each score
Score each domain against the published scale, and record for each score the specific artifacts that support it. Where a domain sits between levels because practice is inconsistent across business units, score the lower level and note the variance rather than averaging; the variance is a finding in its own right.
Keep the scores per domain throughout. If an executive summary needs a single figure, present it alongside the per-domain breakdown, never instead of it.
7. Report the gap, the cause and the cost
A maturity report that lists gaps has done half the work. For each gap between current and target, state what specifically is missing, which other findings share the same root cause, what closing it requires in people, budget and elapsed time, and what the consequence is of leaving it open until the next assessment.
The grouping by root cause is what turns a report into a plan. A dozen findings that all trace back to no asset inventory are one project, not twelve.
Turning maturity levels into a roadmap and a re-assessment cadence
The assessment produces a baseline. The roadmap is what makes it worth having produced, and it is a separate piece of work with its own discipline. Four rules keep it executable:
- Sequence by dependency before priority. Several domains cannot improve until another one does. Access reviews depend on an accurate identity inventory; detection maturity depends on logging coverage; vendor risk maturity depends on knowing which vendors exist. Order the work so the prerequisites land first, even when a downstream item scores worse.
- Move one level at a time, per domain. A plan that takes a domain from level 1 to level 4 in a single year is a plan to write documentation faster than anyone can adopt the practice it describes, and documentation describing practice nobody performs scores worse at re-assessment than an honest level 2. A single-level move per domain per cycle is a plan people can execute.
- Attach each item to a named owner and a funding decision. A gap with no owner and no budget line is a gap that will appear in the next assessment with the same wording. Where an item is deliberately not being funded, record that as an accepted risk with a review date rather than leaving it on the roadmap as perpetual future work.
- Re-measure the domains you funded, when the work lands. Full re-assessment annually, targeted re-scoring of active domains as their remediation completes, and an unscheduled re-score of any domain touched by a significant incident or a major architectural change.
The value of a second assessment lies in its comparability, so re-score against the same model, the same scale and the same scope boundary. If any of the three has to change, report both the old and the new basis for the first cycle after the change.
The items that persist across cycles are worth a second look. A gap that survives two assessments is worth treating as accumulated cybersecurity debt with no owner rather than as a gap in knowledge, and it needs a different kind of intervention than another line on the roadmap.
What makes a maturity score meaningless
Four failure modes turn a scored assessment into a document nobody acts on. Each is worth checking for before the report is circulated.
- Self-scoring with no evidence requirement. When domain owners score their own domains from memory, the scores measure confidence rather than capability, and there is nothing in the record for the next assessor to re-check.
- A single organization-wide number. Averaging across domains hides the variance that matters. A program at level 4 in endpoint security and level 1 in third-party risk does not behave like a level 2 program; it behaves like a program with one open door.
- A target of level 5 everywhere. Uniform maximum maturity is expensive, slow, and on the C2M2's own reading not necessarily optimal. It also makes the roadmap hard to defend to anyone holding the budget.
- A scale that changes between assessments. Re-scoring against a revised model or a re-drawn scope produces a comparison that is not one. Improvement that exists only because the ruler moved is difficult to detect afterwards and expensive to discover in front of a board.
Conclusion
A security maturity assessment is a measuring instrument, and like any instrument it is only as useful as the decision it was calibrated for. Run one because a specific event has put a question on the table that a scored, evidenced, per-domain picture of your program would answer. Score against a model built for that purpose, define the scale before you measure, set targets before you find results, and support every level with an artifact rather than an opinion.
Then treat the output as a baseline rather than a verdict. Sequence the roadmap by dependency, move one level at a time in the domains where the level actually matters, give every item an owner and a funding decision, and re-measure against the same ruler when the work lands. That discipline, repeated, is what a mature security program looks like from the inside.
Do You Know Where Your Security Program Actually Stands?
BD Emerson offers professional Cyber Security Transformation Services that evaluate your security and cyber risk posture, pinpoint the gaps that could be exploited by cyber threats, and offer recommendations for fortifying your defenses. Contact us today to start with a cybersecurity gap assessment.
