For the last three years, the dominant enterprise AI story has been adoption. How many tools, how fast, how broadly. The question companies were asking was not whether to spend on AI but how quickly to get everyone access to it.

That question has changed. Reporting from CNBC and TechCrunch this week confirmed what anyone watching enterprise finance teams has been sensing for months: companies are now actively reducing their AI token consumption. Not abandoning AI. Not reversing their tool decisions. But deliberately cutting back on how much they are spending on tokens, and who is allowed to spend them.

The shorthand for this is tokenminimizing. And it matters, because it signals a shift in how companies think about AI spend. Not as a line item that grows with ambition, but as a cost to be managed with the same discipline as cloud infrastructure or headcount.

Why it is happening now

The answer is visibility. Or rather, the sudden arrival of it.

Gartner published on 24 June that 6% of organisations are already paying more than $2,000 per developer per month in AI token charges, and that AI coding costs will exceed the average developer's salary by 2028.

Source: Gartner, June 2026. Covered by Computer Weekly, The Register, and CIO.

The driver Gartner names is agentic workflows. A single agentic coding task can trigger between 5 and 30 separate model calls that the developer never sees. The bill accumulates in the background. By the time it surfaces in a finance review, weeks of usage have already settled into the invoice.

What changed is that the bill got big enough to surface in conversations that were not originally about AI. And once leadership starts asking questions, the absence of answers becomes its own problem.

What unchecked spend looks like

On the same day Gartner published its analysis, a San Francisco fintech confirmed publicly that a single employee had spent $81,267 on AI coding tools in one week. Not the whole team. One person. One week.

Real example

$81,267 spent by one employee in one week on AI coding. The company had no visibility into individual token consumption until the invoice arrived. This is what the absence of a diagnostic looks like in practice.

This is not a story about reckless behaviour. It is a story about agentic workflows running without cost controls. The employee was almost certainly doing legitimate work. The problem is that the infrastructure around them had no mechanism to surface what it was costing until after the fact.

The Uber response

Uber's answer to the same underlying problem was a hard cap of $1,500 per employee per month on AI tool spend. Not a guideline. A cap. Any employee whose usage exceeds that threshold is cut off until the following month.

Real example

Uber implemented a $1,500 monthly per-employee cap on AI tool spending. The message to employees: AI access is a resource, not a utility. It has a limit, and the limit is enforced.

Uber is not an AI-sceptic company. The cap is not about distrust of AI. It is about the absence of visibility into what the spend was producing, workflow by workflow. When you cannot see what your AI budget is buying, a cap is the bluntest available instrument. It is also the most predictable response.

The problem with tokenminimizing as a strategy

Cutting spend without understanding it is not optimisation. It is rationing. Uber's cap will reduce cost. It will also reduce some genuinely valuable AI usage alongside the waste, because a cap cannot distinguish between the two.

The teams that are navigating this well are not capping first and asking questions later. They are running the diagnostic first. Understanding which workflows justify frontier-model costs, which are running on the wrong model tier, and which have no business running on AI at all. Then making targeted reductions based on what the data shows.

That is a fundamentally different outcome than a blanket cap. One produces savings. The other produces savings and damage.

What to do before the cap arrives

The tokenminimizing trend is not going to reverse. As AI API spend grows as a share of operating costs, the pressure to manage it with the same rigour as cloud or headcount will only increase. The question is whether your team has the data to manage it intelligently, or whether you are waiting for a blunt instrument to land.

A diagnostic that shows you spend by model, by task type, and by team gives leadership something they can act on rather than something they have to guess at. It is the difference between a routing decision and a rationing decision.

The companies that run the diagnostic now are the ones that avoid the cap later.

Know what your AI budget is buying

TokenomicsIQ produces a structured report in 30 minutes. Export your usage data, upload a CSV, get a CFO-readable breakdown by model, task type, and workflow. $3,500 fixed fee. No subscription.

Request early access