AI Unit Economics, FinOps & Infrastructure Cost ModelingPlaybook3 min readUpdated September 2026

Setting Hard Spend Caps So an AI Agent Can't Run Away with Your Bill

An AI agent that can call tools, call itself, or retry on failure has a failure mode a simple chat feature doesn't: it can loop. A looping agent doesn't announce itself, it just keeps calling the model, and by the time anyone notices, the invoice already reflects it.

Token budgeting and circuit breakers are the two controls that keep that failure mode from becoming a finance problem instead of just an engineering bug.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Why agents fail differently than a simple chat feature

A basic chat feature makes one call per user message, so the worst case cost is bounded by how many messages a person can reasonably send. An agent that plans, calls tools, and evaluates its own output can call the model dozens of times for a single user request, and a bug in its stopping logic can turn that into thousands of calls before anything else in the system notices something is wrong.

Setting a per-task token budget

Give every agent task a maximum token budget before it starts, not just a maximum number of steps. Step limits alone can still allow runaway cost if each step involves a large context window, so budget in tokens, and have the agent stop and report back once it hits the ceiling rather than silently continuing. A reasonable starting budget is whatever your most expensive successful task has historically used, with headroom, not an arbitrary round number picked without looking at real usage.

Separate the per-task budget from a per-user or per-account daily and monthly ceiling. A single task can be within budget and still be one of many that, taken together, blow past what a user or account should reasonably cost in a day, especially if something is triggering the same task repeatedly rather than looping within one run.

Building the circuit breaker itself

A circuit breaker needs three things: a hard ceiling on spend or calls per task, a kill switch that actually stops execution rather than just logging a warning, and an alert that reaches a person, not just a log file nobody's watching. The kill switch is the part teams skip most often, building elaborate monitoring that correctly detects the runaway loop and then does nothing to stop it, which defeats the purpose.

A working circuit breaker setup includes these controls:

  • A hard ceiling on spend or calls per task, backed by a per-task token budget rather than a step limit alone.
  • A kill switch that actually stops execution, not one that only logs a warning.
  • An alert that reaches a person, not just a log file nobody is watching.
  • A graceful failure with a clear message when the breaker trips, instead of letting the agent keep trying in a degraded mode.
  • Enough logging at the point of tripping to show what the agent was doing, how many calls it made and what triggered the stop.
  • Tracking of tasks that land close to the ceiling, since they often signal a prompt or tool change that made the agent less efficient.

Deciding what happens when the breaker trips

Trip the breaker and the task fails gracefully with a clear message, rather than trip it and let the agent keep trying in a degraded mode, which often just delays the same runaway pattern. Log enough detail at the point of tripping, what the agent was doing, how many calls it made, what triggered the stop, that an engineer can diagnose the root cause without having to reproduce a live incident from scratch.

Watching the pattern, not just the ceiling

A well-tuned budget rarely trips, which means the interesting signal is tasks that get close to the ceiling without tripping it, since that's often an early warning of a prompt or tool change that's making the agent less efficient before it becomes an outright runaway. Track how often tasks land in the top band of their budget the same way you'd watch a burn multiple drift the wrong way before it becomes a real problem.

Review that near-miss rate on the same cadence you review spend itself, since a rising trend there is the leading indicator, and the invoice is the lagging one. By the time a budget is tripping regularly, the underlying inefficiency has usually already been costing money for weeks.

A runaway shape worth designing against specifically

Say an agent is built to research a topic, draft a summary, then critique its own draft and revise until the critique step is satisfied. That self-critique loop is exactly the shape that produces the worst runaway cases, because the stopping condition depends on the model's own judgment of whether the output is good enough, and a model that's mildly miscalibrated can keep deciding its own work needs one more pass indefinitely.

Design the exit condition for a self-critique loop as a hard step count first, with the quality judgment as a secondary factor, not the reverse. A loop that stops after a fixed number of revision passes, keeps the best one, and ships is bounded no matter what the model decides about its own output. A loop that only stops once the model is satisfied has no such bound built in, and needs the budget and circuit breaker doing all the protective work with nothing structural helping it.

The same failure shape shows up in retry logic: an agent that retries a failed tool call is fine, but one that retries and re-plans from scratch on every failure can rebuild an increasingly expensive plan several times over for what was originally one task. Cap retries at the task level, not just at the individual call level, so a chain of failures doesn't multiply cost at every layer.

Executive Capability Standard

What Good Looks Like

Good looks like every agent task having a real, tested spend ceiling that fails the task gracefully rather than one that exists only in a design doc.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand your agent's actual step and tool-call pattern on a normal successful task before setting any limit on it.
2. Do Manually:Review agent run logs by hand weekly to spot tasks that used unusually high token counts, even if none of them tripped a limit yet.
3. Delegate:Have the engineer who built the agent own the budget tuning, with finance reviewing spend trend rather than setting the technical ceiling.
4. Automate:Build the token budget and kill switch directly into the agent's orchestration code so the check runs on every task without manual review.
5. Buy:Look at an LLM gateway or spend-control product if you're running agents across multiple teams and want a shared enforcement layer instead of rebuilding it per team.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Navan

Putting AI infrastructure spend on a corporate card with a set limit through Navan gives you a second, independent backstop below the software-level circuit breaker, so a runaway agent still hits a real ceiling even if the code-level control fails.

Visit Navan→

Frequently Asked Questions

How do we set the right budget without being overly restrictive?

Start from your actual historical usage on completed tasks, not a guess. Look at the token count for your most expensive normal, successful task and set the budget with some headroom above that, then tighten or loosen it as you see real data on how often tasks hit the ceiling.

Should every user get the same spend cap?

Not necessarily. A cap that makes sense for a trial user exploring the product is often too tight for a paying customer running a legitimate heavy workload, so tiering the cap by plan or account type usually works better than one number for everyone.

What's the difference between a circuit breaker and a rate limit?

A rate limit caps how often calls happen; a circuit breaker caps total spend or call count for a specific task or time window and stops the process entirely once hit, regardless of how quickly the calls arrived. Agents need both, since a slow but unbounded loop can still run up a large bill without ever tripping a rate limit.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides