In this article:

Scoring Your Organization Against NIST AI RMF

Technology
/
July 24, 2026
Scoring Your Organization Against NIST AI RMF

The NIST AI Risk Management Framework does not come with a score, which is why most organizations read it once and file it. You can build the score yourself. Take the four functions, Govern, Map, Measure, and Manage, rate each subcategory you have adopted on a 0 to 4 maturity scale, average within function, and you get four numbers you can defend and re-run next quarter. A first pass for an organization that has never formalized AI governance usually lands between 0.8 and 1.6 out of 4, with Govern highest and Measure lowest. The value is not the number. It is that the number tells you which function to work on next, and in what order.

What you are actually scoring

AI RMF 1.0 organizes into four functions and 19 categories, expanded into subcategories in the companion Playbook. Govern is cross-cutting: policy, accountability, culture, workforce competence, and third party management. Map is context setting: what the system is for, who it affects, what could go wrong, and whether it should be built at all. Measure is analysis: metrics, testing, evaluation, and tracking of the risks Map identified. Manage is action: prioritization, treatment, response, recovery, and communication.

Two things follow from that structure. Govern applies to the whole organization while the other three apply per system, so a single blended score hides a great deal and you should score Map, Measure, and Manage per system before averaging. And Measure depends on Map, because you cannot evaluate risks nobody articulated. That dependency is why Measure scores lowest in almost every first assessment.

The maturity scale

Use five levels. Zero means absent. One means ad hoc, done once by one person. Two means repeatable, a documented practice followed inconsistently. Three means defined and operating, owned by a named role, producing artifacts on a stated cadence. Four means measured and improving, where the practice generates data that changes how you operate.

Score only the subcategories you have deliberately adopted. For the rest, record the decision not to adopt and the reason. The framework is explicitly voluntary and use-case dependent, so a defensible score reports both the rating and the scope you chose. Skipping that step is what makes a self-assessment unusable in front of a customer, an insurer, or a board committee.

Govern: what evidence counts

Look for an AI policy that names decision rights rather than restating principles. An inventory covering internal builds, vendor features, and tools bought on expense cards. Role definitions that place accountability with a person rather than a committee. Third party terms that address training data use, retention, and notification when a model changes underneath you. Workforce training tied to role rather than a single all-hands session.

Evidence that earns a three: the inventory with a last-reviewed date and named owners, the intake form with completed submissions attached, minutes showing a use case was declined, vendor agreements with AI-specific clauses, training completion records. Evidence that does not count: a policy with no approvals recorded against it, an architecture diagram, or a slide describing intended process. Most organizations score Govern between 1.5 and 2.5, because policy is easy to write and decision records are not.

Map: what evidence counts

Map asks whether you understood the context before building. Evidence includes a completed use-case description with intended purpose and explicitly out-of-scope uses, an identified set of affected people including those who are not users, a risk enumeration specific to that system rather than a copied list, and a recorded go or no-go decision with reasoning.

The strongest artifact here is an impact assessment per system, completed before development and updated at material change. Organizations with mature privacy programs usually have the muscle already, because a data protection impact assessment and an AI impact assessment share most of their structure and often the same reviewers. Typical Map scores run 1.0 to 2.0, and the gap is almost always identical: risks were discussed in a meeting and never written down, so nothing downstream can reference them.

Measure: what evidence counts

Measure is the function that separates a governance program from a governance document. Evidence includes defined metrics per identified risk, a test method with results attached, evaluation before deployment and monitoring after it, a record of who ran each evaluation and when, and a method for tracking the qualitative risks that resist metrics.

Concretely: an evaluation set with expected outputs held under version control, accuracy and refusal rates measured against it, adversarial testing results for prompt injection and data leakage, drift monitoring with a threshold that triggers human review, and a feedback channel where users report bad output that someone triages on a schedule. Scores here commonly land between 0.5 and 1.5. The usual failure is thorough pre-launch testing with nothing running afterward, which is a one on this scale regardless of how good the launch testing was.

Manage: what evidence counts

Manage covers what you do about the risks you measured. Evidence includes prioritization traceable to the Map output, documented treatment decisions including risks you chose to accept, incident procedures that name AI-specific scenarios, a tested mechanism to disable or roll back a system, and communication paths to affected people.

The artifacts worth asking for are specific: a risk register entry with an owner and a due date, an accepted-risk memo with a signature, a runbook for reverting to a prior model version or turning a feature off, and at least one recorded incident or near miss with its resolution. Manage scores tend to track Measure within half a point, because you cannot treat what you never measured.

What counts as evidence, in short

  • Artifacts with dates and named owners, not templates
  • Records of decisions, including decisions to decline a use case or accept a risk
  • Results rather than procedures: evaluation output, not an evaluation plan
  • Recurring output on a stated cadence, which is the whole difference between a two and a three
  • Proof a control changed something: a use case narrowed, a vendor rejected, a model rolled back

How the score maps to ISO 42001

AI RMF is voluntary and not certifiable. ISO 42001 is a management system standard you can be audited and certified against. The two align well enough that an RMF score is a reasonable head start on a certification gap assessment, provided you treat the mapping as directional rather than a formal crosswalk.

Govern maps to the management system clauses covering leadership, policy, roles, planning, and competence, plus the Annex A controls on AI policy and internal organization. Map maps to the AI system impact assessment control and the lifecycle and data controls. Measure maps to performance evaluation, monitoring, and internal audit. Manage maps to operational planning and control, plus nonconformity and corrective action. What ISO adds is machinery the RMF never requires: a documented scope, an internal audit program, management review, and a continual improvement loop.

As a rough calibration, organizations scoring 2.5 or higher across all four functions are usually 6 to 12 months from a certification audit. Those below 1.5 are 12 to 24 months out, mostly because the audit trail has to exist for long enough to be sampled. We lay out the sequence in our guide to the ISO 42001 certification path, and ISO 42001 vs NIST AI RMF covers where the two diverge on intent and obligation.

Using the score to sequence work

Do not raise all four functions evenly. The dependency runs Govern to Map to Measure to Manage, and work done out of order gets redone. If Govern is below 1.5, start there with inventory, decision rights, and intake, because nothing downstream holds without them. If Govern is above 2 and Map is below 1.5, standardize one impact assessment template and apply it to your three highest-exposure systems rather than the whole inventory.

If Map is solid and Measure is low, that is the highest-return quarter available to you, since evaluation harnesses and monitoring get reused across every later use case. If Measure sits at 2 and Manage at 1, the fix is small and procedural: a register with owners, a rollback runbook, and an incident scenario added to your existing tabletop exercise.

Rescore every two quarters using the same evidence rule and the same scope. A well-run program moves 0.5 to 1.0 points per function per year. Anything faster usually means the second scorer was more generous, not that the program improved, which is why the evidence list should be kept with the score.

Where BD Emerson fits

We score AI RMF as part of an AI readiness assessment or as a standalone governance baseline, typically in three to five weeks depending on how many systems are in scope. The deliverable is a rating per function with the evidence behind each number, a record of the subcategories you chose not to adopt and why, and a sequenced plan with owners. For organizations heading toward certification, the same evidence feeds directly into AI governance work and an ISO 42001 gap assessment, which is most of the argument for scoring against a framework someone else already reconciled to the standards.

About the author

Leslie Sakal is a Managing Director at BD Emerson focused on cybersecurity, enterprise risk management, and regulatory compliance. She brings over a decade of experience advising organizations across technology, financial services, education, and other regulated industries on implementing organization-wide goals and programs that align with their broader business objectives.
Leslie Sakal
Leslie Sakal
Managing Director