What is AI spending and why is it growing so fast?
AI spending covers all costs a company incurs from AI: API costs from providers like OpenAI, Anthropic, and Google, AI SaaS subscriptions, and the infrastructure costs of running AI workloads. It grows faster than expected for three reasons: usage compounds non-linearly as products gain adoption; most teams lack workflow-level visibility so inefficiencies accumulate undetected; and the token subsidy that kept frontier model prices artificially low during 2022-2025 is ending. UBS research from June 2026 found that 60% of companies actively managing AI budgets have already moved to cheaper models in response to cost pressure. Full analysis here.
What is AI token cost?
AI token cost, also called LLM token cost, is the per-token price charged by AI API providers for model usage. Providers bill separately for input tokens (the text you send, including your system prompt and conversation history) and output tokens (the text the model generates). The token invoice is one of nine cost buckets in a typical AI deployment. The other eight do not appear on any vendor bill, which is why teams that manage only what they see on the invoice are usually working with an incomplete picture. Full breakdown here.
What are the nine AI cost buckets?
Identified at FinOps X in June 2026: (1) Token Invoice, the only one on any vendor bill; (2) Infrastructure overhead; (3) Orchestration cost; (4) Prompt processing waste, including system prompt bloat; (5) Caching inefficiency; (6) Agent loop redundancy; (7) Model-task mismatch; (8) Sync-on-batch overhead; (9) Uncached context. TokenomicsIQ analyses all nine in every diagnostic and ranks recommendations by estimated monthly saving.
What file formats do you accept?
OpenRouter CSV exports, Anthropic Console usage exports, OpenAI dashboard exports, and Typing Mind JSON conversation exports. Analysis quality is highest with OpenRouter data, which includes per-request token counts, named application context, and caching data. Any limitations in the source data are noted transparently in the report methodology section.
What does the report actually contain?
Six sections: a personalised executive summary, spend breakdown by model and named workflow, inefficiency analysis ranked by monthly saving, specific recommendations with engineering implementation notes, risk flags covering vendor concentration and model lifecycle exposure, and a methodology section. Both PDF and HTML formats included.
How long does it take?
Under 30 minutes from upload to download for most datasets. Larger exports (over 100,000 rows) may take up to 45 minutes.
Is my data stored?
No. Your usage export is processed entirely in memory and deleted once your report is downloaded. We do not store, index, or analyse client data beyond the duration of the report generation process.
What if I do not use OpenRouter?
The tool works with Anthropic Console and OpenAI dashboard exports. The analysis is slightly less granular for these sources because their standard exports do not include per-request prompt content or application-level attribution. We note this transparently in the report. Typing Mind exports provide rich conversational context and work well for teams using AI assistants across client workflows.
Can I run a follow-up report after implementing the recommendations?
Yes. The delta report ($1,250) runs the same analysis on your new usage data alongside your original export and shows exactly what has changed: saving confirmed, remaining opportunities, and any new patterns since the baseline. You re-upload both CSVs when you are ready. We never hold your data between sessions.