AI
Insights
True Velocity: Measuring the Real Impact of AI on Software Delivery

We track delivery performance across more than 300 projects through an internal platform we call Pulse. We built it because understanding whether AI is actually improving delivery requires looking beyond activity alone.Ask how much faster AI has made teams, and you will get very different answers depending on who you ask. Vendors building and selling AI tools naturally focus on the biggest potential gains, while others responding to the hype focus on the limitations, because ambitious claims deserve scrutiny. Neither perspective tells the full story, because both often measure what is easiest to measure rather than what actually matters.
That is the challenge with productivity. It is not one thing. We think about this in terms of true velocity: not simply how quickly teams produce outputs, but how quickly they can move from an idea to a valuable, reliable outcome.
More code written, more pull requests merged and more output produced in a day are all measurable improvements, but they are not the same as delivering more business value. A tool can make a task dramatically faster without changing whether the outcome is more useful, more reliable or more valuable. The pattern emerging from research and industry experience is similar: time saved creating code is often reallocated to reviewing and validating it. AI can increase throughput, but the bottleneck does not disappear. It moves to where judgement is required.What we measure instead
Rather than focusing only on output volume, we look at measures such as lead time, rework, first-time-pass rate and defects. These give a better picture of the health of the delivery process itself, not just how much is produced.
What this shows is that AI can create meaningful improvements, but those gains are not evenly distributed. They depend on the type of work, the maturity of the process and the level of risk involved. It also means being careful about simple claims, since there are too many variables involved to confidently say that one factor alone caused a specific improvement. A single productivity number covering everything from initial idea to production release may sound attractive, but it rarely reflects the complexity of building real software.
Two different economics
This is also why comparisons around AI productivity can become misleading. The cost of testing an idea has fallen dramatically, while the cost of embedding that idea safely inside a live, running business has changed far less. Those are two different economics.
It has never been cheaper to test whether an idea has potential. A working version can often be created faster than the research process that would once have recommended building it. That creates a genuine opportunity to build, test and learn earlier. But the value only comes from understanding which parts of the process have become cheaper and which still require expertise, engineering and judgement.
Measuring what matters
That leads to a few practical principles.
First, productivity improvements in experimentation should not automatically be treated as productivity improvements in production delivery, since testing an idea and safely operating a product are different challenges.
Second, teams should separate exploration, delivery and assurance rather than applying one productivity assumption across all three. A credible view of AI impact should recognise that some activities become significantly faster, while others remain dependent on expertise and careful decision-making.
Third, delivery plans should include assurance plans. For anything with real consequences, it is not enough to explain how quickly something can be built. Teams also need to explain how it will be tested, monitored, released safely and supported when something goes wrong.
Finally, success should be measured by outcomes rather than activity. Features shipped and hours spent are easy to count. The harder and more meaningful measures are whether the product works, whether users adopt it and whether it delivers the intended result.
The bigger picture
Generating a plausible answer - a draft, a feature or a working version of a product - is becoming cheaper. Knowing whether that answer is good enough to trust, and designing the process to find that out safely and quickly, is where the value has moved.
AI has not removed the work that makes software valuable. It has changed where that work happens: earlier into context and problem framing, later into validation and assurance and increasingly into the judgement calls that decide what is worth building.
True velocity is not about producing more. It is about creating better outcomes, with the confidence that what reaches users is worth trusting. That is the work worth organising a project, a team or a delivery partnership around, and it is where we focus our time.







