Somewhere this quarter, a finance director is looking at an AI bill that’s roughly double what it was three months ago, and nobody in the building can tell them why.
Not because the team is careless. Because the thing driving the bill, which model handled which task, how many tokens a given agent burned working through a document, how often a workflow retried itself before landing an answer, isn’t visible to anyone until the invoice arrives.
Nobody owns that number in real time. Everybody inherits it after the fact.
This is where a lot of enterprise AI investment sits right now. Gartner’s own research on generative AI adoption found that 54% of firms require demonstrable ROI to justify broad investment, and only 22% currently report getting significant value from their GenAI tools.
That gap isn’t mainly a capability problem. The models work. It’s a visibility problem. You can’t prove ROI on spend you can’t see broken down by task, by team, or by model.
Token Costs vs. Licence Costs: Why Traditional Budgeting Fails
Token costs behave nothing like the licence costs enterprises are used to budgeting for:
- RPA Licences: A flat, predictable, annual number.
- Token Bills: Dynamic spend that moves with usage, model choice, and how well or badly a workflow is built.
Route a routine task to a frontier model when a smaller one would do the job, and the cost difference on that single call might be small. Do it across a few hundred thousand calls a month, invisibly, because nobody’s routing logic is checking, and it stops being small.
Gartner flagged a version of this directly in its Hype Cycle work on computer use agents: for high-frequency tasks, token-based automation is currently more expensive than a traditional annual RPA subscription, and the variability makes budgeting for it genuinely difficult.
That’s not a niche edge case. That’s the shape of the problem most enterprises are about to run into at scale.
Token Sprawl: The New Bot Sprawl
It’s worth naming what this actually is, because it’s not new. It’s the RPA licence sprawl problem again, one layer further up the stack:
- Bot Sprawl: Paying for capacity that sat idle behind static schedules while queues backed up elsewhere.
- Token Sprawl: Paying for model calls that go to the wrong engine for the job, with no one watching which agent, which process, or which team is driving the spend, and the cost only becoming visible once it’s already been incurred.
Same failure mode. Different unit of waste.
Enter Agentic FinOps: The Market Gap
Gartner has already named the category forming to answer this: agent management platforms, offering what it calls agentic FinOps, tracking cost, token usage, and utilisation across agent deployments to control sprawl and prove ROI.
Gartner rates the category high in benefit and still emerging in maturity, and it’s honest about the gap:
- Cost control, rate limiting, and workload optimisation mechanisms aren’t fully mature yet.
- Budgets for the governance tooling to fix this are lagging behind the desire for it.
In other words: everyone can see the problem, but almost nobody has built a mature answer to it.
The Need for Independent Visibility
What has to exist is a layer that can see cost and performance across every model and every agent a business runs, not a single vendor’s usage dashboard for its own platform, in the same way a genuine orchestration layer has to see across every automation platform rather than just one.
It’s the same architectural argument applied to a different kind of spend. We’ve spent the last few years teaching enterprises to ask who’s watching their bot estate and what it actually costs to run. The same question now applies to every token their agents spend, and most organisations don’t yet have anyone in the room who can answer it with a straight face.
This is exactly the direction our own work on orchestration and governance is heading next, extending the same independent, vendor-neutral visibility we’ve built for RPA licence and utilisation economics up into model and token spend. More on that shortly.
For now, the practical starting point is simpler: before the next AI budget review, find out who in your organisation can currently tell you, with a straight face, what last month’s token spend actually bought.
Take Control of Your Token Spend Live in London
Uncontrolled token consumption and agent sprawl are fast becoming Q1 board-level issues. Join us at C TWO Connect on 15 October 2026 in central London to learn how leading enterprises are getting ahead of AI economics.
Hear directly from automation and AI leaders at Lloyds Banking Group, Commerzbank, and E.ON on how they measure real-time cost and attribution across bots and agents, without relying on vendor-owned dashboards marking their own homework.
Free to attend for end users of automation and AI technology. Places are limited.