Build Your Own AI Inference Cost Model in Three Tabs
Most AI cost surprises trace back to the same root cause: nobody modeled the spend before it happened. A support team turns on a new model tier, a product feature starts calling the API on every page load, and finance finds out when the cloud bill lands.
A working cost model does not need to be complicated. It needs three things: the assumptions that drive spend, the arithmetic that turns those assumptions into a number, and a place to test what happens when they change. Here is how to build that in a spreadsheet you actually keep open.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What actually belongs in the model
Before you open a spreadsheet, list the handful of inputs that move your bill. For most teams that is:
- Tokens in and tokens out per request, measured separately, since output tokens usually cost more
- The per-model rate you're paying, by tier, since a fallback to a cheaper model changes the math
- Your cache hit rate, since a cached response costs a fraction of a fresh one
- Requests per active user per day, and how that scales with usage
Everything else in the model is just this list multiplied out. If you can't state these four numbers for your own product right now, that's the gap to close first, not the spreadsheet layout.
Structuring three tabs instead of one
A single sprawling tab is where these models go to die. Split it instead:
- An assumptions tab holding every rate, ratio and price, sourced and dated, so you know when it went stale
- A calculation tab that references the assumptions tab and does the multiplication, one row per model or feature
- A scenario tab that lets you flex volume, cache rate or model mix and see the total move without touching the other two tabs
Keeping calculations separate from assumptions means you can hand the assumptions tab to an engineer to keep current, without engineering ever touching your formulas.
Turning tokens into a cost per customer
The number that actually matters to a CFO isn't total token spend, it's cost per active customer, because that's what you compare against what the customer pays you. Say your product sends an average of 40 requests per user per day, each costing a fraction of a cent after your cache hit rate is applied. Multiply that by days in the month and you have a per user cost you can set next to your average revenue per user. If that ratio is moving the wrong way as usage grows, you've found the problem before the invoice does.
Where these models usually go wrong
The failures are consistent across companies:
- Modeling only the primary model and ignoring fallback or retry calls, which can double effective usage during an outage
- Leaving out embedding calls, moderation calls and any background job that also hits the API
- Using a cache hit rate from launch week instead of the current, usually lower, rate as usage diversifies
- Never revisiting the model after the first build, so it quietly stops matching reality
Reviewing the model every month, not once
Treat this the way you'd treat a burn multiple review: a monthly check, not a one time build. Pull the actual invoice, compare it to what the model predicted, and adjust the assumptions tab for the gap. Companies that keep spend efficient tend to watch their burn multiple stay in a healthy band even as usage grows, and an AI inference cost model is one of the few tools that lets you see that trend before it shows up in the income statement.
Stress testing the model against a pricing or model change
A cost model that only reflects today's rates is one vendor announcement away from being wrong. Providers change per-token pricing with little notice, and engineering teams routinely swap which model tier handles a given feature once a cheaper or better option ships. Use the scenario tab to run that change through the model before it happens in production, not after the invoice shows the impact.
Say you're deciding whether to move a feature from a premium model tier to a cheaper one for routine requests, keeping the premium tier only for cases that fail a quality check first. Change the model mix assumption in the scenario tab and watch the blended cost per customer move, holding request volume and cache hit rate constant so you're isolating the one variable that actually changed. If the swap holds up in a side by side quality test, you now have a real number to bring to the team deciding whether to make the change, instead of a guess that the cheaper tier should save money.
Run the same exercise whenever a provider announces a rate change, up or down, before the new rate takes effect. A model that only gets touched after a change lands in the invoice is reactive; one stress tested ahead of the change is what lets finance flag a cost shift before it happens rather than explain it after.
What Good Looks Like
Good looks like a spreadsheet that predicts next month's AI invoice within a small margin, updated from real usage data rather than launch week guesses.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
How often should I update the assumptions tab?
Update it whenever a rate changes, a new model tier goes live, or actual invoices start drifting from the model by more than a few percent. Most teams find monthly is the right cadence even without a trigger, since usage patterns shift gradually and a stale assumption is easy to miss otherwise.
Can finance build this without engineering's help?
Finance can build and own the spreadsheet, but needs engineering to supply the real numbers: actual requests per user, real cache hit rates, and which model tier handles which feature. Ask for those four numbers once, then own the model yourself and only go back to engineering when something structural changes.
What's the biggest blind spot in a token cost model?
Background and retry calls. Teams model the obvious user facing request and forget embeddings, moderation checks, and automatic retries after a timeout, which can add a meaningful share of total spend without ever appearing in a product spec.
About the numbers
This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.
Related Guides
Edge vs Cloud AI Inference: When On-Device Actually Pays Off
How to find your own crossover point between on-device AI inference and a cloud API, once you count hardware, model limits, and update infrastructure.
Capturing the Batch API Discount Without Hurting Your Product
How to find the AI calls that can tolerate a delay, move them to a discounted batch endpoint, and recover real margin without touching real-time features.
Does Routing Inference Across Regions Actually Save Money?
When routing AI inference requests to whichever region is cheapest actually pays off, and the latency and complexity costs that can erase the savings.
What Evaluating Your AI Agent Actually Costs to Run
See where AI agent evaluation cost comes from: judge-model calls, human review and test set upkeep, with a worked run example and ways to keep spend in check.
Why Faster AI Responses Cost More, and When to Pay for It
How batching, model size, and dedicated capacity trade off against each other, and a simple way to decide which of your AI features actually needs to be fast.
The Real Cost per Resolved Ticket Once AI Handles Support
Why cost per ticket looks better with AI support automation than it actually is, and how to calculate a number that accounts for escalations and rework.