While reviewing our delivery metrics, we identified a number that appeared to demonstrate exceptional delivery velocity: 98% of pull requests merged within 24 hours.
On the surface, it suggested a high-performing engineering system. Looking deeper, we realised the metric measured speed but not the things that mattered just as much at scale: review quality, approval discipline, and production confidence. Our measurement framework had matured faster than our governance model around it.
So we made a different decision. We held the number back, strengthened the review and approval process around it, and kept the original figure for what it was genuinely useful for: a baseline to measure the mature process against. The lesson has shaped how we think about AI adoption today. A metric that looks better without governance than with it is not measuring performance. It is measuring the absence of an operating model.
That experience frames how we read the wider landscape. Across the organisations we work with, the first wave of AI adoption is largely complete: AI is drafting content, summarising meetings, assisting developers, answering internal questions, and automating parts of existing workflows. The harder question is what comes next, because experimentation is moving faster than governance, operating models, measurement frameworks, and workflows are adapting. Organisations are adopting AI the way many organisations historically built their digital platforms: fast to launch, without the foundations required to scale.
This is not an argument against AI. We use it daily and the productivity gains are real. It is an argument about what happens after the initial enthusiasm, because buying AI tools does not create transformation. Sustainable value comes from redesigning how work gets done, and we hold that view because we have had to implement it ourselves.
The question we are now seeing organisations wrestle with is not whether to use AI, but how to turn AI usage into a repeatable, measurable capability.
The difference between using AI and building AI capability
An AI operating model is the structure around the tools: who owns AI decisions, what information can enter which systems, how output is reviewed, how workflows change, and how outcomes are measured. It is the difference between simply using AI and being changed by it.
In our view, AI maturity is not measured by the number of tools an organisation has deployed. It is measured across four capabilities.
AI fluency. Not tool training or prompt technique, but the ability of teams to identify where AI creates genuine value, apply judgement to its output, and incorporate it responsibly into their work. We define this more precisely below, because we had to measure it in ourselves first.
Workflow redesign. AI creates limited value when it is added on top of existing processes. The organisations that benefit most redesign how work moves: which steps disappear, which decisions change, who owns the output, and how the capacity created gets used.
Governance. Governance determines whether AI can be trusted in operational environments that matter: ownership, data boundaries, review standards, and accountability.
Measurement. Measurement determines whether AI investment translates into business outcomes. Without it, organisations can see activity but not impact.
These four capabilities separate organisations experimenting with AI from organisations building AI capability. The rest of this article works through each, using our own implementation as the evidence.
What implementing AI across our delivery model taught us
Before AI became something we advised organisations on, it became something we needed to operationalise ourselves. As we scaled AI across our delivery model, we ran an evidence-based capability review using our internal delivery systems to understand where our adoption actually stood.
The finding was clear: AI experimentation had become part of our culture and how our team worked, but the next stage was moving from individual productivity gains to repeatable, attributable workflows. Tools and automation existed; the opportunity was creating the operating discipline to make those improvements measurable and scalable.
The review changed how we operate. AI-assisted workflows are now assigned to named individuals, with clear ownership and visible human judgement throughout the process. Adoption that cannot be understood cannot be improved.
It also gave us a practical definition of AI fluency: an organisational capability built around five behaviours identifying where AI creates genuine value, applying judgement to its output, redesigning workflows around it, governing what it produces, and measuring outcomes. Fluency is not demonstrated by tool usage; it is demonstrated through repeatable workflows, accountable ownership, and evidence of improved outcomes.
A simple test is to examine any AI-assisted process and ask three questions: who owns it, what outcome has it improved, and where was its output last reviewed by a person? The quality of those answers reveals whether AI has become an operating capability or remains a collection of individual experiments.
The rest of our early journey will be familiar to any organisation adopting significant new technology. Workflows were refined through several iterations before they held up in real delivery conditions. Some automations that performed well in testing needed redesign when exposed to real client requirements. Review standards evolved as we learned where AI output required stronger human judgement. None of this was unusual; it is the normal maturity curve of adopting a new capability. The technology was rarely the hardest part. The hardest part was changing how work gets done.
Activity shows usage. Outcomes show transformation.
A pattern we see frequently: organisations measuring AI adoption and interpreting it as progress. Licences purchased, tools deployed, prompts created, tokens consumed, active users. These metrics are easy to collect because they measure activity, and they answer a procurement question rather than a business one. They tell an organisation that AI is being used. They do not tell it whether the business is better.
The harder metrics are the ones that matter: time genuinely returned to teams, cycle times reduced, quality maintained or improved, operational friction removed, capacity created, and revenue enabled by AI-supported work.
When we built our own measurement approach, the first decision was defining what we would not measure. Lines of code were excluded, because activity volume is not a meaningful indicator of engineering quality. Raw AI activity was treated as context, never performance. Metrics we could not source reliably were marked unavailable rather than estimated. The purpose of measurement is not to create impressive numbers. It is to create better decisions, which is why our scorecard was designed to produce a worklist: the specific factors influencing each outcome and the actions required to improve them. When we introduced additional quality assurance stages, they were measurable from their first day, because the process was designed with measurement in mind. That is the discipline AI adoption requires: decide what would fool you, then design it out.
We apply the same standard to our own AI journey. We are building measurement infrastructure alongside adoption, which means we are establishing baselines rather than attempting to recreate them afterwards. Any organisation claiming precise AI ROI without understanding its pre-adoption baseline is making an assumption, whether it intends to or not. The question that should be asked is not how much AI the organisation is using. It is what measurable business outcome has improved because of it.
The measurement challenge is not new; we see the same pattern across digital platforms. Our audits regularly identify organisations where valuable activity exists but cannot be understood or acted on. In one engagement, content represented more than 90% of traffic with no conversion measurement configured. In another, a discovery crawl revealed a client had almost three times the product catalogue they believed existed. The technology was not the problem. The missing capability was measurement, and AI adoption without measurement repeats the pattern at greater speed and greater cost.
Exploration is a phase, not a strategy
None of this means early experimentation is wrong. It is how organisations learn, and it is how we learned.
The exploration phase involves testing use cases, understanding limitations, discovering where AI creates genuine value and where it creates polished but unreliable output, understanding true costs, and establishing initial governance. Organisations should expect this phase, resource it, and treat it as learning rather than failure.
The optimisation phase is where value compounds: scaling proven use cases, redesigning workflows around them, embedding AI into operating models rather than alongside them, assigning ownership, and measuring return. This is not a timeline, and organisations move at different speeds depending on their systems, people, and risk environment. The risk is not experimentation. The risk is remaining there indefinitely: endless pilots, expanding tool inventories, and no mechanism for turning lessons into repeatable capability.
External research points the same way. MIT’s GenAI Divide research identifies organisational integration and learning, rather than technology capability, as the major barrier between AI experimentation and measurable business impact. The technology works. The operating models around it often lag.
Governance is an accelerator, not a brake
A second pattern we observe: AI decisions without clear ownership. Who approves a new tool? What information can enter which systems? Who reviews AI-generated output, and who is accountable when something goes wrong? Where those answers are unclear, complexity accumulates quietly. Technology without ownership creates risk.
Our own governance was built through implementation and continues to evolve. An early lesson from automating at scale was identifying overlapping deployment paths into the same production environment. The solution was not another tool. It was a governance decision built into the operating model: one canonical repository, controlled production deployment, and documented rules around configuration changes. The principle applies directly to AI: the question is not whether an automation can act, but who decided it should.
The same thinking shaped how we introduced AI into our support workflow. We invested more design effort in what surrounds the model than in the model itself. Data boundaries were defined before workflows went live, not discovered after. Tickets are tiered across four complexity levels, which determine how much AI assistance is appropriate for each scenario. AI output does not move into staging or production without named human approval. And each AI run is logged with its cost, so the price of a workflow is visible run by run rather than hidden inside a licence fee. The model is only one component. The surrounding operating system is where the value is created.
Our standing principle is simple: treat AI output as a second reviewer, not a rubber stamp. Governance is not what slows AI down; it is what creates the confidence to use AI in work that matters. One example from our own quality process demonstrates why: a pre-launch review identified AI development tooling still connected within a build and flagged it before release. AI accelerated the development. The operating model made it safe to ship.
When everyone can create content, know what only you can know
Generative AI has dramatically reduced the time and cost of producing content, and the consequence is already visible: more content, published faster, sounding increasingly similar. When everyone can create volume, volume itself becomes less valuable.
What becomes scarce is what AI cannot create independently: first-hand experience, original data, lessons from real delivery, specialist expertise, and trust built over time. The organisations that stand out will not be the ones producing the most content. They will be the ones producing content that only they could produce. AI should accelerate expertise, not replace it.
The same principle is reshaping discovery. As people increasingly use AI assistants alongside traditional search, visibility depends less on publishing volume and more on becoming the source worth referencing: demonstrable expertise, structured information, authority signals, trusted references, and communities that validate your knowledge. Content now has two audiences, the people consuming information and the systems helping them find it, and both increasingly reward the same thing: useful expertise, clearly structured.
AI does not create capacity. Redesign does.
The quietest pattern we observe is often the most expensive. AI is frequently added on top of existing processes: same workflows, same roles, same deadlines, plus new tools to learn. Individual tasks become faster, but the organisation does not, because the surrounding system has not changed. Teams end up maintaining the previous way of working while learning the new one, which is not efficiency gained. It is complexity added.
We recognise this from our own early implementation, and it is why we now treat workflow redesign as the unit of AI value. The questions are which steps disappear, which decisions change, who owns the output, and what happens with the capacity created. The question we would put to any AI initiative is simple: is this creating capacity, or adding another layer of complexity?
Before scaling AI, ask five questions
For leaders moving from experimentation to optimisation, the discipline is not another tool evaluation. It is five questions.
- What business outcome are we trying to improve? Not which tool to deploy. What gets better, for whom, and by how much.
- How will we measure success? Outcomes, not activity, against a baseline established before scaling. Decide what would fool you, then design it out.
- Who owns AI governance? A name, not a committee. Clear data boundaries, review standards, and accountability.
- Which workflows need redesign? AI layered onto an unchanged process returns only a fraction of its potential.
- Do teams have the capacity and capability to adopt this well? This is the fluency question: ownership assigned to named people, judgement applied to output, and time allocated for redesign.
These five questions are the practical expression of the four capabilities: answering them well is what fluency, redesign, governance, and measurement look like in operation. They apply equally whether an organisation has five people or five thousand. The scale changes; the discipline does not.
Organisations that can answer these questions are positioned to scale. Organisations that cannot are still exploring, and there is no issue with that when exploration is deliberate. Exploration is progress. Scaling without answers is spending.
The takeaway
AI transformation is not won by organisations that acquire the most tools. It is won by organisations that redesign work, govern change, and measure outcomes. Our own experience reinforced the lesson: a fast-looking metric taught us that activity without context creates the wrong conclusion, and the operating model we built afterwards proved more valuable than the metric itself.
The next advantage will not come from access to AI, because access is becoming universal. It will come from turning AI into a repeatable capability: owned by people, embedded into workflows, governed appropriately, and measured against outcomes.
For organisations building digital platforms, the same principle applies whether the capability is a website, a content system, or an AI workflow: the technology matters, but the operating model determines whether it creates lasting value.
AI & workflow
WordPress
Business & growth