How to Manage AI Budgets in the Age of Tokenmaxxing
In the third and final session of FinOps Summer Camp, we covered how to turn AI spend visibility into action.

The question this session answered: once you can see your AI spend, what do you do about it? Rem Baumann, Director of Product at Vantage, and Augusto Leal, Sales Engineer, laid out a framework for governing AI cost and setting budgets on it.
Tokenmaxxing was never going to last
Managing AI spend is a top priority right now because Tokenmaxxing is running out of road. It started as a flex: through 2024 into 2026, teams treated token consumption as a proxy for productivity and built leaderboards around it. It's a vanity metric. More usage doesn't mean more output, and the invoices made that obvious.
None of this is new. The same thing happened in early cloud migrations, when teams measured success by workloads shipped rather than money saved. This time the correction came faster. Uber put four times more engineers on frontier AI tools this year and still drove cost per token down by treating efficiency as an engineering problem, not a restriction on who gets access. That's the shift: from rewarding usage to rewarding the people who get the most value out of their tokens.
Governing AI spend comes down to four steps: visibility, allocation, measuring value, and budgets.
1. Visibility: see the full breadth and depth of AI spend
AI spend now spans categories that used to sit in separate buckets: application AI (models embedded in your own product), enterprise AI (the coding tools and assistants employees use daily), and self-hosted inference on providers like Modal or Baseten, billed by compute instead of by request.
Most of what you need to allocate that spend is hidden in the bill. A provider invoice tells you what you paid, not which customer, feature, or team drove the cost. Vantage pulls billing and telemetry into one view so the number carries those attributes with it. Without that, you have a total and no way to act on it.
2. Allocation: attribute every dollar to a team
Allocation is what makes a cost mean something. Some costs map cleanly. If a chatbot serves customers, you can tag each request and trace spend back to the customer or feature behind it.
Shared costs are harder. One API key might cover three teams; one block of inference might support the whole product. For those, proxy allocation splits the cost by a metric that approximates real usage, so every dollar still gets mapped.
Vantage handles this with virtual tags, a middleware layer that applies tags in bulk from business logic instead of asking developers to label every resource by hand. The result is full coverage for showback or chargeback, normalized across providers so a model or user reads the same whether the spend came from Anthropic, OpenAI, or Cursor.
3. Measure value: make sure your tokens are worth it
Once spend is visible and attributed, you can ask whether it was worth it, and no single metric answers that. Any one number can be gamed: shrink your pull requests and tokens-per-PR drops while cost-per-feature stays flat; reuse prompts to pad your cache-hit ratio while the bill climbs anyway. The answer is a spread of value-based KPIs read together: cost per developer, model mix, cache-hit ratio, and output measures like cost per customer or tokens per resolved ticket.
To make the metrics meaningful, roll them up to something the CFO cares about. One developer running hot on token usage is noise; a whole organization trending over budget is a signal. Watch the trend over time and use it to show engineers which usage is paying off.
4. Budget: set limits without slowing teams down
Budgets come last because they depend on everything before them. Once you know the value you're getting, you can set a number, company-wide down to a single developer, and decide what happens when someone hits it. That decision matters more than it sounds: downgrading to a cheaper model, throttling, or stopping usage each carry different consequences, and the right choice depends on who feels the pain.
For internal tools, avoid hard stops. An engineer who can't ship because a budget was capped will find a way around it. It's trickier when the AI is customer-facing. Your users feel any downgrade you make, so there's less room to blow the budget in the first place, and the pressure shifts to efficiency work like caching. Run both dollar/token budgets and value-based budgets, and settle the fail-over behavior up front.
What's next
That wraps up FinOps Summer Camp. Catch up on the other sessions if you haven't:
- The State of Cloud and AI Cost Management in 2026
- How Make.com Built AI Cost Observability That Engineers Actually Use
If you want the visibility, allocation, and budget views from this session running on your own AI spend, that's the core of what Vantage does.
The team's around on LinkedIn and in the Vantage Slack community and FinOps Foundation Slack community if you want to keep the conversation going.
Sign up for a free trial.
Get started with tracking your cloud costs.

