Bank of America is calling it a shift from tokenmaxxing to token rationing. BCA Research, writing in the FT this week, calls it the end of the token subsidy. For three years, AI model makers priced their products well below cost to drive adoption. That era is ending, usage-based pricing is now the direction of travel, and companies that have not mapped their spend at workflow level are walking into the new model blind.

What tokenmaxxing looked like

During the subsidy era, the implicit message from AI providers was: use as much as you want. Flat-rate plans and aggressively below-cost API pricing created an environment where teams had little incentive to think carefully about token efficiency. More tokens meant richer context, better outputs, and faster iteration. The cost was someone else's problem, covered by investor capital flowing through model maker balance sheets.

The result was predictable. Teams built workflows on frontier models for tasks those models were never required for. System prompts accumulated instructions over months without review. Caching was skipped because the cost difference was not visible enough to justify the engineering effort. Whole engineering functions were running on AI budgets that nobody had actually approved.

"The end of the token subsidy marks a seismic shift in how AI spend is priced, controlled, and justified inside organisations." — BCA Research, via FT Unhedged, July 2026

What token rationing looks like

GitHub moved from flat fees to usage-based pricing in April 2026. It was not the last. As model makers come under pressure to stop subsidising adoption and start generating returns, usage-based pricing is becoming the norm. Every token now has a real and visible price attached to it.

Token rationing is what happens when that price lands on a finance team's desk. It looks like Uber capping engineer AI spend after burning through its annual AI budget in four months. It looks like a San Francisco fintech discovering one employee spent $81,267 in a single week. It looks like a CFO asking a question that the engineering team cannot answer with confidence: what are we actually spending this on?

In practice, token rationing means companies are making active decisions about which models to use for which tasks, setting spend limits, and in many cases switching to cheaper open-source alternatives for lower-stakes workflows. UBS research published in June 2026 found 60% of companies actively managing AI budgets have already made this move.

Why the teams without visibility are most at risk

The transition from tokenmaxxing to token rationing is straightforward for teams that already have workflow-level visibility into their AI spend. They know which tasks are running on which models, what each workflow costs per month, and where the optimisation opportunities are. They can make the routing change with confidence and quantify the saving before and after.

The teams without that visibility face a harder problem. They can see the total monthly bill but not what is driving it. They know roughly which providers they are using but not which internal workflows are consuming the most spend. When pressure comes from finance to reduce AI costs, their only tool is blunt: cut the spend cap and hope nothing breaks.

That approach is both riskier and less effective than a structured diagnostic. Cutting the wrong workflow causes quality problems. Leaving the actual waste in place means the saving is smaller than it should be. And it leaves the finance conversation unanswered: you have reduced the number, but you still cannot explain what the spend was actually for.

The diagnostic comes before the rationing decision

The companies navigating this transition well are not rationing first and asking questions later. They are running the diagnostic first: mapping spend at workflow level, identifying model-task mismatches, and making routing decisions based on actual data. The saving is larger, the change is safer, and the finance conversation becomes a capability demonstration rather than a damage-limitation exercise.

Token rationing without visibility is an arbitrary constraint. Token rationing with workflow-level data is a deliberate cost-control strategy. The difference between those two outcomes is whether you ran the diagnostic before or after the pressure arrived.

The era of the token subsidy is ending. The teams that enter the new model with their spend already mapped are in the stronger position.