The organisations that engineer for AI cost reality now will define the economics of enterprise AI for the rest of the decade.
The fastest-growing AI cost optimisation activity inside most enterprises right now is procurement negotiation. Discount terms, committed-use rebates, multi-vendor pressure plays. It is the natural reflex of an organisation watching its bill climb.
It will not work. Negotiating discounts on subsidised prices is not cost optimisation. It is reducing the rate at which value is being destroyed. When OpenAI is losing $1.22 for every dollar of revenue, and Anthropic has publicly admitted that its subscription pricing was not built for production agent workloads, the discount you negotiate today is an artefact of a pricing structure the market has already decided cannot continue.
Real AI cost engineering starts somewhere else - inside your own architecture and operating model. In Part 1 of this series, we set out the macro-economic shift, the end of the token subsidy era, and the three patterns that make AI cost rise even as token prices fall: the headline-price illusion, frontier magnetism, and the context tax. This piece is the prescription.
Gartner predicts that more than 40 per cent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls. MIT's Project NANDA found that 95 per cent of generative AI pilots are failing to deliver measurable P&L impact. The FinOps Foundation found 73 per cent of organisations exceeded their AI budget projections in the past year. The pattern is consistent: the technology is delivering on capability and falling short on economics. Gartner's own research identifies the inverse pattern.
Organisations with successful AI initiatives invest up to four times more in data and platform foundations than their peers. The capability gap is not where most boards assume it is. It sits in the unglamorous discipline of measurement, architecture, and governance, applied to AI cost the way it has historically been applied to cloud cost. In our engagements across financial services, manufacturing, and the public sector, four disciplines consistently separate AI programmes that scale from those that quietly bleed margin.
In the experimentation phase, model selection was a procurement decision driven by the lowest published price per token. In the third act, it is an architecture decision driven by cost per business outcome. This means matching the model to the task, not the task to the model.
Routine, high-frequency, low-complexity workloads belong on smaller, more efficient models. Frontier reasoning models are reserved for high-margin, genuinely complex work where the additional token consumption is justified by the value of the answer. Gartner's senior analysts have made the same case publicly: indiscriminate use of frontier reasoning for tasks that do not require it is one of the largest sources of silent token inflation.
You cannot optimise what you cannot see. The first task in AI cost discipline is to make consumption visible at the level the business cares about: cost per use case, cost per workflow, cost per outcome. This is the AI extension of a discipline that mature cloud teams already practice. FinOps for cloud has been around for years; FinOps for AI is its natural next chapter, and it is moving fast.
The FinOps Foundation's 2026 report identified AI as the fastest-growing new spend category among its members. The disciplines transfer cleanly: tagging, attribution, showback and chargeback, anomaly detection, predictive forecasting. What is different is the velocity - AI consumption can spike by an order of magnitude in a single sprint when a developer enables an agentic loop.
Each of the three patterns identified in Part 1 has a corresponding engineering response. For the headline-price illusion, run controlled, vendor-agnostic cost tests on representative workloads whenever a new model is being considered. Do not accept the price card as evidence of cost. For frontier magnetism, build a model gateway that makes model selection a policy decision rather than an individual preference.
Routine workloads go to smaller models by default; access to frontier requires justification. For the context tax, redesign agent workflows to minimise re-sent context. Published research has shown that disciplined context engineering alone can reduce per-task token consumption by 60 per cent or more without measurable loss of accuracy.
The history of every previous platform rationalisation - from on-premise to cloud, from monolith to microservices - is that the organisations that survived the price changes were the ones with portability. The same will be true for AI. This is not about hedging your bets on which vendor will win. It is about not building business-critical workflows on a single provider's pricing assumptions, when those assumptions are visibly changing every quarter.
Multi-provider architecture is harder to build, but it is the only structural protection against pricing decisions made by a counterparty losing money on every transaction.
Across our client base we are applying the same approach to AI cost that we have applied to enterprise data platform cost for years. The Decision Inc. FinOps practice extends naturally into FinOps for AI: making consumption visible, attributing it to the workloads that drive it, and engineering the architecture so that cost moves predictably with value rather than independently of it.
For organisations earlier in the journey, our AI Cost Diagnostic establishes a current-state baseline of model usage, consumption patterns, and exposure to provider pricing changes, then defines the architectural and operating-model changes needed to bring cost under control.
The diagnostic is led by certified data and AI architects, scoped to deliver insight in weeks rather than months, and designed to feed directly into focused implementation sprints with quantified outcomes.
The subsidy era taught organisations to consume AI without measuring it. The third act will reward the opposite discipline. The organisations that engineer for that reality now will define the economics of enterprise AI for the rest of the decade.
The organisations that wait for the pricing to settle will discover that the pricing has been settled against them. Book an AI Cost Diagnostic to benchmark your organisation's exposure and define the architectural disciplines needed to survive the rationalisation.
Talk to our Data & AI team to benchmark your organisation's exposure and define the architectural disciplines needed to bring AI cost under control.