a
Executive Perspective • AI & Advanced Analytics

The Third Act of
Enterprise AI:

When the token subsidy ends

Henri De Bruine By Henri De Bruine
Director: Data Analytics, Decision Inc. UK 2026
Executive Perspective Series: The Third Act of Enterprise AI  ·  Part 1 of 3
483%
Increase in average enterprise AI spend, 2024 to 2026
Gartner 2026
73%
Of organisations exceeded their AI budget projections
FinOps Foundation 2026
85%
Of enterprise AI budget is now inference, up from 40% in 2023
Analytics Week 2026

Have you noticed recently that your AI assistant subscription does not take you quite as far as it used to? You ask the same questions, run the same actions, and two documents into the process: "You have reached your usage limit."

There is a shift underway in Generative AI platform economics. If you are watching the enterprise AI market as closely as we are, a pattern emerges. We are moving through three distinct stages, and most organisations are about to feel the third one in their budgets.

Tokens were cheap, credits were generous, and nobody was watching the meter. As providers pivot from land-grab to profit, the unit economics of every deployed use case will be re-examined, and the assumptions baked into a lot of business cases will be tested and probably broken.


The Three Stages

Moving Through Three Distinct Stages
Most Organisations Will Feel the Third

Stage one was experimentation. The question was simply "can it do this?" Proofs of concept, pilots, a thousand demos. Cost barely entered the conversation, because the goal was to prove capability, not to run anything at scale.

Stage two is production. This is where most serious organisations started moving from impressive demos to embedded, dependable workflows. The hard, unglamorous work of making AI actually load-bearing, and value generating, inside the business.

Stage three is the rationalisation of token economics, and it is coming fast. For the last few years, the major model providers have priced aggressively to win adoption. Much of what we have been consuming has been, in effect, subsidised. That cannot last.

As providers pivot from land-grab to profit, the unit economics of every deployed use case will be re-examined, and the assumptions baked into a lot of business cases will be tested and probably broken.

The subsidy was never going to hold
article_third_act_subsidy.jpg
The Subsidy

The Subsidy Was Never Going to Hold

The scale of the subsidy is now visible in public filings and analyst reports. In Q1 2026, OpenAI reported an adjusted operating margin of negative 122 per cent. The company loses $1.22 for every dollar of revenue it brings in.

SemiAnalysis has calculated that a heavy user on a $200-per-month Claude Max plan can deliver a gross margin of minus 900 per cent for the provider. The flat-rate subscriptions and aggressive API prices that built the market were always loss leaders, designed to seed adoption while market share, not margin, was the goal.

The correction has already started. In May 2026, Anthropic split its subscription billing into two pools, separating chat from agent workloads. The head of Claude Code put it plainly: the subscriptions "weren't built for the usage patterns of these third-party tools." GitHub moved Copilot to consumption-based pricing in June.

OpenAI is, according to the Wall Street Journal, weighing further pricing changes ahead of its IPO.

The cost pressure on enterprise customers is already visible. Uber's CTO publicly confirmed the company had exhausted its entire 2026 AI budget by April. JP Morgan analysts published an internal note this year titled simply: "AI Bills Are Out of Control."

The Silicon Valley shorthand for this moment is "token maxxing": consuming large volumes of tokens, often without a clear return.

At Decision Inc. we are already seeing this play out. Across multiple client engagements, optimising the cost of AI solutions has moved from a future concern to a priority activity. And it is happening for a reason most organisations have not yet seen.


The Paradox

The Price Per Token is Falling.
The Cost Per Answer is Rising.

The cost per answer is rising
article_third_act_parachute.jpg

The most uncomfortable dynamic of the current market is not the eventual end of the subsidy. It is what is happening before the subsidy even ends.

Two numbers from the analyst community look impossible to reconcile. Gartner forecasts that by 2030, the cost of running inference on a one-trillion-parameter model will drop by more than 90 per cent compared to 2025. Epoch AI's analysis of frontier model benchmarks shows per-token prices have already fallen between 9x and 900x per year for various performance milestones. By any normal economic logic, enterprise AI bills should be shrinking.

They are doing the opposite.

Average enterprise AI spend rose from $1.2 million in 2024 to $7 million in 2026, an increase of 483 per cent. Gartner's worldwide forecast puts global AI spending at $2.59 trillion in 2026, up 47 per cent year on year. The FinOps Foundation's 2026 State of FinOps report found that 73 per cent of respondents saw their AI costs exceed their original budget projections.


The Three Patterns

Token prices are collapsing. Bills are climbing.
Three Patterns Explain Why.

In our work with clients across financial services, manufacturing, and the public sector, three patterns recur. Once you can name them, you can start to manage them.

Pattern 1: The Headline-Price Illusion

A vendor reduces the published price per token from one model generation to the next. Procurement logs a saving on the price list. The invoice goes up.

The reason: the newer model is almost always a reasoning model, and reasoning models consume multiples of the token count to answer the same question. Gartner's March 2026 analysis found that agentic and reasoning models require between 5 and 30 times more tokens per task than a standard generative AI chatbot.

Bain & Company's analysis identifies the same effect: model prices are falling roughly 10x per generation, but the effective cost per task is staying flat or rising, because reasoning consumes the savings before they reach the customer. We have been running our own controlled tests at Decision Inc. The same questions, with the same context, fed into different models from the same vendor, comparing older generations against newer reasoning-oriented ones. The pattern is consistent: a headline price cut on the model card does not survive contact with real workloads. What looks like a procurement win on the price list is often a cost increase on the bill.

Pattern 2: Frontier Magnetism

When the next model ships, no one stays on the cheaper, older generation. They upgrade. Bain calls this the frontier magnetism problem: last generation gets cheaper, but everyone moves to frontier, where prices stay high.

The savings on the older model are real but unclaimed, because no team is willing to defend deploying it in production once a more capable version exists. The compounding effect is significant. Frontier models cost more per token than the model they replace, consume more tokens per task because of their reasoning behaviour, and are typically applied to use cases that did not previously exist. Each of those three forces pushes the bill in the same direction.

There is a Gartner warning that captures this exactly: "Chief Product Officers should not confuse the deflation of commodity tokens with the democratization of frontier reasoning." The cheap tokens are the ones you have stopped using.

Frontier Magnetism
article_third_act_pattern2.jpg

Pattern 3: The Context Tax

Agentic workflows are not one model call. They are dozens. The Stanford Digital Economy Lab has measured that re-sent context accounts for 62 per cent of total agent inference bills.

The agent calls a tool, reads the result, decides what to do next, calls another tool. Each step requires the model to re-process the entire conversation history before generating the next action. Most of what an enterprise is paying for in agentic AI is the model re-reading what it already knows.

This is the single biggest reason that pilot economics never match production economics. The pilot was scoped as a single chatbot query. The production deployment is a multi-step agentic loop running thousands of times per day.

The ROI calculation that justified the build assumed chatbot-level token consumption per workflow. The real number is an order of magnitude higher.


The Root Cause

The Root Cause

The root cause of AI cost inflation
article_third_act_root.jpg

These three patterns share a common cause: the unit of measurement that matters to the business has shifted, and most cost models have not shifted with it.

The relevant unit is no longer cost per token. It is cost per answered question, cost per resolved support case, cost per processed invoice, cost per closed lead. The headline price-per-token metric, which dominated procurement conversations in the experimentation phase, is now actively misleading. Two models can have identical price-per-token numbers and a 5x difference in real cost per business outcome.

This is not a future problem. Analytics Week's 2026 Inference Economics report found that AI inference now accounts for 85 per cent of the enterprise AI budget, up from 40 per cent in 2023. The cost has already shifted from training to running. The conversation has not.

The Implication

What Happens When the Subsidy Ends

The most uncomfortable implication of these three patterns is that the silent cost inflation is already happening, before the public price increases land. Most enterprises have a token economics problem they cannot see, because they are measuring the wrong unit.

When provider pricing rationalises, organisations that have not built measurement and right-sizing discipline will not face one cost shock. They will face two: the loss of the subsidy, and the discovery that their consumption per outcome was already trending in the wrong direction.

The experimentation phase rewarded curiosity. The production phase rewarded discipline. The third phase will reward those who treat token economics as seriously as they treat the technology itself.

The price per token is falling. The cost per answer is rising. The organisations that survive this gap will be the ones that learn to measure, and manage, both.

In Part 2 of this series, we set out the prescription: the four disciplines that consistently separate AI programmes that scale from those that quietly bleed margin.

The Close

AI economics are not a procurement problem.
They are an architecture problem.

The organisations that engineer for that reality now will define the economics of enterprise AI for the rest of the decade. The organisations that wait for the pricing to settle will discover that the pricing has been settled against them.

Decision Inc. is an advisory-led cloud & AI partner that works with over 300 organisations annually to build and scale enterprise data platforms across financial services, retail, telecoms, manufacturing and the public sector.

Decision Inc. Data & AI

Book a Diagnostic Session

Benchmark your organisation's exposure to the rationalisation of AI token economics and define the architectural disciplines needed to survive it.