adnenaSend a note
Insights

Artificial Intelligence

Scaling Beyond AI Pilot through Six-Move Capability Cycle

Scaling Beyond AI Pilot through Six-Move Capability Cycle

Why pilots succeed, why scaling breaks them, and what the Six-move Capability Cycle does differently

On 27 February 2024, Klarna and OpenAI announced that its AI assistant had handled 2.3 million conversations in a month; the equivalent, Klarna said, of 700 agents, cutting resolution time from eleven minutes to under two. It was treated as a triumph.

Fifteen months later, chief executive told Bloomberg the push had gone too far: quality had slipped, and Klarna was hiring humans back. By its Q3 2025 earnings call, the assistant was doing more work than ever, the equivalent of 853 agents, saving $60 million, while a human tier had been deliberately rebuilt around the interactions that still needed one. Same company, two verdicts eighteen months apart. What changed was not the technology. It was whether anyone was running a loop underneath the headline number.

That distinction is what separates organisations still running pilots from ones compounding an advantage a rival cannot simply copy. Here is why pilots so reliably succeed, why that success becomes a liability at scale, and what the six-move Capability Cycle does that a pilot alone cannot.

The Pilot’s Advantage

A pilot is built, almost always unintentionally, to succeed. It runs on curated data by people who understand it, a small self-selected team more tolerant of friction than the wider organisation will be, an expert quietly catching the model’s mistakes before anyone else sees them, a success metric the pilot team chose itself, and a sponsor senior enough to clear away procurement and legacy obstacles that would stall an ordinary project for months.

Every one of those features removes variance, which is exactly why a pilot’s results are real but not yet representative. RAND Corporation’s 2024 study of 65 experienced data scientists and engineers found more than 80% of AI projects fail overall, roughly twice the rate of non-AI IT projects and that 84% of practitioners who lived through a failure blamed leadership and organisational decisions ahead of the data or the model itself. Pilots rarely expose that gap, because a pilot is precisely the phase when leadership attention and organisational goodwill are at their peak.

Where the Pilot Meets the Real Organisation

Scaling removes every one of those conditions at once. The curated data becomes the organisation’s full, messy record; the expert safety net disappears, because the point of scaling is to remove the need for one; the self-selected enthusiasts are replaced by an ordinary, sometimes careless, workforce and customer base.

Two examples show what follows.

Air Canada found the same gap in a small-claims tribunal: its chatbot wrongly told a grieving customer he could claim a bereavement discount retroactively, and when Air Canada argued the bot’s own words shouldn’t bind the company that published them, a tribunal disagreed in February 2024 and ordered the fare refunded. The sum was trivial; what it exposed was not - nobody at Air Canada was monitoring what the chatbot told customers against the airline’s own, regularly updated policy. Neither organisation had skipped testing. Neither had a mechanism to notice, in production, that reality had diverged from the pilot’s assumptions before it became a write-down or a ruling.

S&P Global’s 2025 enterprise survey found the share of firms abandoning most of their AI initiatives had already jumped from 17 to 42% in a single year. Building the missing mechanism is what the Capability Cycle is for.

**The Capability Cycle: Six Moves
**
**1. Set the hypothesis: **Define what should get better, by how much, before anything is built.

2. Establish the baseline: Measure how the work runs today. Without this, nothing that follows can be proven.

**3. Measure the outcome: **Track what actually happened against the hypothesis - the gap.

**4. Diagnose the failure: **Study where and why it fell short. Most programmes skip this, because it means telling a sponsor the number they announced does not hold at scale.

**5. Redesign the workflow: **Change decision rights, data access or approval steps - not just the tool.

6. Retrain the people: Teach employees to evaluate machine output, then feed what they learn into the next hypothesis, which starts the next cycle.

Moves three and four are where the real learning happens, and it is institutional, not computational.

MIT’s Project NANDA, in July 2025 study, found the generative tools behind most stalled pilots simply could not retain feedback between sessions. An underwriter learns which flagged claims to trust; a chief information officer learns which legacy system will never expose the data a redesign needs and none of that shows up in a vendor’s roadmap.

BCG attributes roughly 70% of realised AI value to people and process and only 10% to the algorithm;

McKinsey’s August 2026 survey found AI ‘high performers’ - 6% of firms, unchanged from a year earlier - are 3.3 times more likely to pursue transformative change, and 75% of them have redesigned a workflow because of it, against a quarter of everyone else.

What Compounds

Klarna’s second act shows the loop working in reverse: by the end of 2025 it was running a larger, more carefully segmented deployment than the one it had to publicly correct.

JPMorganChase shows it at real scale: more than 200,000 employees now use its internal LLM Suite, with over 100 generative-AI tools in production, and chief executive Jamie Dimon told shareholders in April 2026 that AI would “affect virtually every function, application and process” at the bank.

Neither is finished. That is the point: each is several cycles ahead of a rival still running its first pilot.

One Assumption Worth Challenging

Not every scaling failure is a learning failure, and it is worth being honest about the cycle’s limits. RAND’s own data complicates the tidy version of this argument: since 84% of failures trace to leadership, not the model, a meaningful share is a diagnosis that would embarrass whoever championed the pilot and that will not get run honestly just because a mechanism exists on paper. Nor is persistence the same as learning: BCG found 94% of companies plan to keep investing in AI even without near-term returns, which is easy to mistake for the Capability Cycle in action. It is not. Funding a pilot without diagnosing why it has not scaled is a very expensive.

Before the Next Board Update

The organisations leading were never running more pilots. They were running fewer, and extracting more from each before moving on, which asks a sponsor to admit what went wrong before expanding it, and asks the board to fund diagnosis and redesign as seriously as it funds licences and compute. Klarna’s executive did that in public, at real reputational cost, and came out the other side running something larger than what he corrected.

Learning is your IP. The only question worth putting to the next AI update is not how big the next pilot will be. It is what, specifically, was measured, diagnosed and redesigned since the last one.

About the author

Mahesh Tanwani

Founder & CTO, adnena

All insights