In this article:

The AI Readiness Checklist

Technology
/
July 23, 2026
The AI Readiness Checklist

An AI readiness checklist works best as a set of pass or fail tests across five dimensions: data foundations, governance and risk, security posture, skills and operating model, and use-case portfolio. Score each dimension 0 to 3 against evidence someone can produce on request, add the numbers, and the total tells you whether to build, fix, or wait. Most organizations score between 6 and 10 out of 15 on a first honest pass, with data foundations and governance pulling the average down. What follows is the working version: three concrete tests per dimension, what a top score actually looks like, and how to turn the total into a sequence of work rather than a slide.

How to score, and why evidence is the whole trick

Use a 0 to 3 scale per dimension. Zero means the capability does not exist. One means it exists in one team and nowhere else. Two means it exists across the organization but is manual, inconsistent, or undocumented. Three means it is documented, owned by a named person, and produces an artifact you could hand to a new engineer or an auditor without narration.

The evidence rule matters more than the scale. Ask for the catalog entry, the intake form, the access review, the evaluation results. Self-reported readiness scores run 30 to 50 percent higher than evidence-based ones, and that gap is exactly where programs stall at month six. If nobody can produce the artifact in ten minutes, the score is one lower than the room believes. For a fuller treatment of what an assessment evaluates and how the dimensions interact, see our piece on what an AI readiness assessment covers.

Dimension one: data foundations

Three tests. Can you name the system of record for each of the three data domains your first use cases need, and does one team own each? Do you know freshness and completeness for those domains as measured numbers rather than assumptions? Can a new analyst find, access, and understand a table without asking a person?

Ready looks concrete. A catalog covering the top 20 to 50 tables with owners and business definitions. Documented refresh intervals with monitoring that alerts when a pipeline misses. Access granted through groups rather than one-off tickets. Lineage that shows where a field came from. Not ready looks like three competing customer tables, spreadsheets serving as the reconciliation layer, and one analyst who is the documentation.

You do not need a finished data platform to start. You need one domain clean enough that a model output can be trusted, plus a clear statement of which domains are not there yet, so nobody promises a use case that depends on them.

Dimension two: governance and risk

Three tests. Is there an inventory of AI systems in use, including the features your SaaS vendors turned on? Is there a documented approval path for new use cases with named decision rights? Does your risk method produce different answers for different use cases, or the same answer every time?

Ready means an inventory refreshed at least quarterly that covers internal builds, vendor features, and tools people expensed themselves. It means an intake a product manager can complete in under an hour and get a decision on within two weeks. It means your controls map to a recognized structure such as the NIST AI Risk Management Framework or ISO 42001, so you are not reinventing governance per project, and it means a written position on human review for consequential decisions.

In most environments we review, the inventory turns up two to three times as many AI-touching systems as leadership expected, and the majority are vendor features rather than anything engineering built. The usual failure is a policy that requires approval for all AI use with no path to actually get approved, which teaches teams to route around it. Structuring that intake and the control set behind it is the core of AI governance work.

Dimension three: security posture

Three tests. Do you know where prompts and outputs are logged, how long they are kept, and who can read them? Is access to models and AI vendors authenticated through your identity provider with credentials you can revoke in one place? Have you tested at least one AI-specific failure mode against your own architecture rather than reading about it?

Ready means retrieval systems that respect the permissions of the source system instead of querying a flattened index. It means keys in a secret manager with rotation, logging with a deliberate retention decision, and an assessment of the AI risks specific to your design: prompt injection through documents users can upload, data leakage through retrieval, agents holding write access they do not need.

The most expensive readiness failure we see is a retrieval index built by crawling everything a service account can read. It demos well and surfaces compensation data in week three. Fixing that after launch costs 3 to 10 times what scoping permissions correctly would have.

Dimension four: skills and operating model

Three tests. Is there a named owner accountable for AI outcomes, rather than a committee? Has at least one team shipped and then maintained a model or language model feature in production for 90 days or more? Is there a defined answer to who evaluates output quality after launch, and how often?

Ready means product, data, engineering, and security sitting in one working group with a decision cadence rather than a quarterly readout. It means engineers who can read and extend an evaluation harness, a named accountable person per deployed use case, and a budget line for run cost rather than build cost alone.

Run cost is where estimates break. An internal assistant serving a few hundred employees commonly costs $3,000 to $30,000 per month in inference, retrieval infrastructure, and platform fees once usage is real, and that number rarely appears in the pilot business case. If the pilot lives with a contractor or one enthusiastic engineer, and no job description includes maintaining it, score this dimension a one no matter how good the demo was.

Dimension five: use-case portfolio

Three tests. Does each candidate have a written value hypothesis with a baseline number captured before launch? Does the portfolio mix payback speeds, or is it five versions of the same idea? Are there written kill criteria?

Ready means 8 to 15 scored candidates, each rated on value, data readiness, and risk, with two or three selected for the next two quarters. It means a measured baseline for each: hours per week, cost per ticket, cycle time in days, error rate. And it means a documented answer to what result would cause you to stop, agreed before anyone builds.

Baselines are the item teams skip and regret. Without one, a working tool cannot be defended in the next budget cycle, because nobody can say what it changed. Portfolio construction and sequencing is most of what AI strategy work actually produces.

Reading the total

Add the five scores for a total out of 15.

  • 0 to 5. Not ready to build. Pick one domain, one owner, and one narrow internal use case, and spend a quarter on foundations before committing budget to anything customer-facing.
  • 6 to 9. Ready for one narrow use case with a named owner and a hard scope boundary. This is where most mid-market organizations land. Ship one thing, instrument it, and use what breaks to fund the next round of foundation work.
  • 10 to 12. Ready to run a small portfolio in parallel. Two or three use cases, shared evaluation and governance, one platform decision rather than per-team tooling.
  • 13 to 15. Ready to scale. The constraint is now adoption and change management, not capability, and the risk shifts from technical failure to building things nobody uses.

Sequencing the fixes

Fix the lowest-scoring dimension that blocks your first use case, not the lowest score overall. Governance at one does not block an internal document summarizer; data foundations at one does. Security at one blocks anything touching customer data regardless of how strong the other four are. Work the blocking dimension for 60 to 90 days, rescore with the same evidence rule, and only then widen scope.

Two rules save the most time. Do not buy a platform to fix a governance score, because tooling records decisions and cannot make them. And do not start with the use case that has the largest business case if it depends on your worst data domain, because program credibility is set by whether the first thing shipped and stayed shipped.

Where BD Emerson fits

We run this checklist as a structured AI readiness assessment, typically over two to four weeks, with interviews, artifact review, and a scored result your leadership team can argue with. The output is a score per dimension, the evidence behind each number, the items you decided not to pursue, and a sequenced 90 day plan naming owners. Clients who already know where they stand usually skip to the strategy and governance work instead. Either way the first conversation is about which use case matters, because the checklist only means something measured against a specific thing you are trying to ship.

About the author

Leslie Sakal is a Managing Director at BD Emerson focused on cybersecurity, enterprise risk management, and regulatory compliance. She brings over a decade of experience advising organizations across technology, financial services, education, and other regulated industries on implementing organization-wide goals and programs that align with their broader business objectives.
Leslie Sakal
Leslie Sakal
Managing Director