An AI maturity assessment measures how repeatably an organization turns an idea into a production system that stays in production. Platform size is a weak proxy for it, and so is the model count. Deloitte’s State of AI in the Enterprise, a survey of 3,235 leaders in 24 countries run in August and September 2025, shows the gap this opens: 42% of leaders rate their AI strategy as highly prepared while rating themselves lower on infrastructure, data, risk and talent, only 34% use AI to reinvent products or core processes, and one in five has a mature governance model for autonomous agents. Declared maturity has drifted from demonstrated maturity.
What an AI Maturity Assessment Actually Measures
An AI maturity assessment is a structured evaluation of how reliably an organization delivers and operates AI systems, scored across dimensions such as data, technology, governance, people and process. Its output is a ranked list of constraints and the order in which to remove them, not a grade.
What decides whether the exercise is useful is the line between adoption and capability. Adoption counts activity: licences issued, pilots started, teams trained. Capability describes what happens to the next use case, the one nobody has built. A company where 4,000 people hold an assistant licence and a new production system takes nine months of negotiation is high on adoption and early on the AI maturity curve.
Why Most Maturity Assessments Stop at a Score
The systematic review of AI maturity models published in PeerJ Computer Science screened 686 records and analyzed 15 primary studies. Two of its findings explain why so many AI maturity assessment tools end at a number. Six of the 13 studies that assess maturity aspects, 46%, look mainly at technology, because technology is what an outsider can count. Only seven of the 15 models, 47%, were validated. The review also reports that no study covered every critical dimension it identified, and a model missing dimensions will not reflect an organization’s real capability.
None of that makes the category worthless. A maturity model puts a fragmented capability on one page and makes a board conversation specific. It fails when opinion supplies the inputs and the report ends at 2.7 out of 5, a number that says nothing about which meeting to hold on Monday.
The Seven Dimensions Worth Scoring
The same review identified seven critical success factors recurring across 13 of the models examined: data, analytics, technology and tools, intelligent automation, governance, people and organization. An AI maturity assessment worth the time scores each of them from one to five and makes every score cite an artefact. The frame comes from published models, not one vendor’s methodology.
Data maturity here means access under working conditions, not a company-wide cleanup program. Readiness is measured per use case, against the data it needs; maturity is measured per organization, across the capability that delivers use cases. The per-use-case audit sits in our piece on AI-ready data foundations.
Technology and tools is where MLOps maturity sits, and the manual step count between commit and serving traffic is the score; why teams stall there is covered in why enterprise AI pilots stall. Governance scores as an operating capability with an external reference point: the four functions of the NIST AI Risk Management Framework, govern, map, measure and manage. That is the shape we use when setting up AI model governance.
The Five AI Maturity Levels, Described in Delivery Terms
Five levels is a convention rather than a finding. The PeerJ review reports five as the most frequent count while individual models use nine levels or a ten-step ladder, so the scale matters less than what is attached to it. The descriptions below use delivery behavior instead of adjectives, because “experimenting” and “transforming” are hard to disagree with and harder to verify.
- Level 1, Ad hoc. Work starts when someone volunteers. Every project renegotiates access, tooling and approval, so delivery time is unpredictable and a failing model is noticed by a complaining user.
- Level 2, Experimenting. Pilots run and some work, but time to production is unknown because nothing has finished the trip. Deployment escalates to whoever will sign, and monitoring lasts as long as the builder’s interest.
- Level 3, Operationalized. A system runs in production with a named owner and a documented release path. Delivery time is measurable for work resembling what already shipped, new patterns still stall, and a quality alert reaches whoever answers it.
- Level 4, Scaled. A new use case reuses infrastructure, pipelines and controls that exist, so delivery time can be quoted in advance and roughly holds. Release follows a standing rule, and rollback and retraining are routine.
- Level 5, Optimized. The portfolio is managed on value: use cases are retired as well as launched, cost per use case is known, and failures change the standard rather than one project.
Evidence Beats Self-Reporting: How to Run the Assessment
To assess AI maturity, ask four evidence questions per team and per business function, and require an artefact behind every answer. Score the dimensions from what those artefacts show. A 60-item questionnaire yields a number in a day; an AI maturity assessment run on documents yields a diagnosis.
- How long did the last use case take from approval to its first real decision? Evidence: the change record.
- How many use cases run in production today, and who is paged when one breaks? Evidence: the model register and on-call rota.
- Is there one document describing the path from idea to release that teams actually follow? Evidence: the document and the last two use cases that used it.
- What happened the last time model quality dropped? Evidence: who noticed, through which signal, and how many days passed before anyone acted.
Each has a paper trail or visibly does not, which settles a level faster than a long survey. Deloitte’s data shows what self-rating does: strategy scores highest of the dimensions leaders rate, and that question is usually answered by the people who wrote the strategy. Maturity is also uneven, so a fraud team at level 4 and a marketing team at level 1 average to a 2.5 that describes neither.
Reading the Result: The Constraint, Not the Average
The average of seven dimension scores is the least useful number an AI maturity assessment produces, because a program moves at the pace of its weakest dimension. An organization scoring four on six dimensions and one on the seventh does not behave like a 3.6. It behaves like a one, with expensive assets queued behind the constraint.
Three profiles cover most results:
- Strong technology, weak governance. Models get built quickly and stop at a risk review with no defined entry point. The fix is a control inside the pipeline with a named approver, not more platform.
- Strong data, weak organization. Models are accurate and unowned. Nobody holds budget for the year after launch, so systems pass to a team that never asked for them and decay.
- High self-rating, thin artefacts. Every dimension scores three or four and no register, runbook or incident record supports it. This profile measures ambition, and the honest re-score lands two levels lower.
A report worth paying for names the constraint and estimates what removing it costs. The better ones also name the next use case that will prove it has gone. That turns enterprise AI maturity from a slide into a decision.
From Assessment to Roadmap: What Changes at Each Level
The sensible move from any level is to the next one; a level 2 result with a level 5 target for the same year buys platform capacity nobody can fill. From level 1, take one use case to production end to end and write down the route. From level 2, name an owner for that route and pick the first component worth reusing. From level 3, make the next two use cases cheaper than the first.
A level means something only when attached to a next step, which is why vendor adoption paths are sequenced rather than graded. Microsoft’s Cloud Adoption Framework for AI runs through strategy, plan, ready, govern and manage stages, each carrying work rather than a label.
Most roadmaps converge on one target state: repeatable delivery, governance running inside the pipeline, and a portfolio measured on value realized. That is the AI factory operating model, and we build it with clients as an enterprise AI Factory. Our service page for that industrialized AI delivery model puts the problem plainly: most organizations have run AI pilots, and fewer than 30% reach production. Where the constraint sits in mandate rather than tooling, the work belongs to a wider enterprise AI transformation.
An AI maturity assessment earns its cost when it changes what you do next quarter. If you are holding a result nobody agrees on, talk to our AI delivery team and we will run the evidence questions with your leads.
Frequently Asked Questions
What is an AI maturity assessment?
An AI maturity assessment is a structured evaluation of an organization’s ability to deliver and operate AI systems repeatably, scored across data, technology, governance, people and organization. A useful one ends with the constraint that limits delivery and the order of repair, not a single score.
How many levels are there in an AI maturity model?
Usually five, though the count is a design convention rather than a standard. The PeerJ systematic review of 15 AI maturity models found five the most frequent choice, with individual models using nine levels or a ten-step ladder. What each level describes matters more than the count.
What is the difference between AI readiness and AI maturity?
Readiness is measured per use case, against the data and controls that case needs, and it answers whether you can start. Maturity is measured per organization, across the capability that delivers use cases, and it answers whether you can keep delivering. An AI readiness assessment can pass while maturity sits at level 2.
How long does an AI maturity assessment take?
Weeks rather than months when it runs on evidence. Collecting artefacts across a handful of teams takes longer than a survey and far less than a discovery program.
Sources
- Sadiq R. B. et al., Artificial intelligence maturity model: a systematic literature review, PeerJ Computer Science 7:e661, 2021.
- Deloitte AI Institute, State of AI in the Enterprise, 3,235 leaders, 24 countries, 2025.
- NIST, AI Risk Management Framework (AI RMF 1.0).
- Microsoft, Cloud Adoption Framework for AI: AI strategy.


.webp)
