Why AI-Native Software Runs Lower Gross Margins Than SaaS
AI-native software runs lower gross margins than classic SaaS because every customer interaction can trigger a real, variable inference cost, which behaves more like a services business's cost of delivery than a hosting bill. Classic SaaS costs little to serve once the product is built, so its margins have historically been among the strongest of any industry.
That difference shows up directly in gross margin, and it's worth understanding before you benchmark your own numbers against the wrong reference class.
Vendors Covered in this Article
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
What actually sits in cost of goods sold for each model
Classic SaaS cost of goods sold is mostly hosting, a customer success or support allocation, and third-party data costs, all of which scale slowly relative to revenue. An AI-native product adds inference cost that scales roughly with usage, and usage doesn't scale neatly with revenue the way a per-seat SaaS contract does: a customer paying the same subscription fee can drive wildly different inference volume month to month depending on how heavily they use the product.
That volatility is the part a classic SaaS finance model was never built to handle. A per-seat contract lets you forecast cost of goods sold from headcount alone; an AI-native contract needs a usage forecast too, and a customer who ramps up usage faster than expected moves your margin even though the invoice they're paying hasn't changed at all.
Why do AI-native margins run below classic SaaS?
Software companies broadly post some of the strongest gross margins of any industry1, and that figure is the reference point most board members and investors still carry in their heads when they look at your gross margin line. An AI-native product running meaningfully below that reference isn't necessarily unhealthy, it may simply be a different cost structure, but you need to be the one explaining that gap before someone else assumes it's a problem.
The cost decision classic SaaS never had to make this often
A traditional SaaS company's main margin decision is infrastructure efficiency: moving to cheaper compute, optimizing a database query. An AI-native company has that same decision plus a second one that matters more: which model tier handles which request, since routing a simple query to an expensive frontier model is pure margin erosion with no quality benefit the customer would ever notice.
Caching is the third decision that doesn't map cleanly onto the classic SaaS playbook: a cached response costs a fraction of a fresh model call, and a product where many requests are similar enough to cache well can improve its margin substantially without touching the model or the infrastructure at all. Products where every request is genuinely unique don't get that option, which is itself a useful thing to know about your own cost structure.
How do you present a lower gross margin to a board?
Show gross margin trending in the right direction as volume grows, since inference cost per customer should improve with better routing, caching, and negotiated vendor rates even if the absolute margin percentage stays below a classic SaaS benchmark. A board that sees the trajectory understands the story; a board that only sees a single quarter's number next to a generic software benchmark will ask why you're behind, without the context of what's actually driving it.
When you show gross margin to a board or investors, cover these points:
- Show the trend as volume grows, since routing, caching and negotiated vendor rates should lower inference cost per customer over time.
- Explain that a margin below the classic SaaS reference is not automatically a red flag, and point to trajectory rather than one quarter.
- Report inference cost as its own line so readers can see it behaves differently from ordinary hosting cost.
- Confirm inference spend is coded consistently as cost of goods sold, so the number you present is actually accurate.
Getting the accounting right so the number is even accurate
None of this benchmarking means anything if inference cost is miscoded as an operating expense instead of cost of goods sold, or the other way around, which happens more often than it should when a vendor bill gets routed through whatever code the finance team last used for a similar-looking invoice. A bill-pay platform like BILL is where that coding decision actually gets made line by line, and it's worth a specific policy for AI vendor invoices rather than leaving it to whoever processes the bill that month.
Write the policy down once, which model providers' invoices count as cost of goods sold, which internal tooling spend counts as opex, and apply it consistently across every new vendor rather than deciding case by case, since an inconsistent policy makes quarter-over-quarter margin comparisons meaningless even if each individual quarter's number is defensible on its own.
What Good Looks Like
The standard is a gross margin figure with AI inference cost clearly broken out, trending in a direction you can explain, rather than a single number compared once against a benchmark built for a different cost structure.
Building The Capability (5-Stage Skill Ladder)
How to Get Started
Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.
Frequently Asked Questions
Is a lower gross margin than classic SaaS automatically a bad sign?
No, but it does mean the classic SaaS benchmark isn't the right yardstick on its own. Look at the trend as much as the level: a margin that's improving as you get better at routing and caching tells a healthier story than a flat number compared once against an industry figure that assumes a different cost structure entirely.
Should we separate AI inference cost from other hosting cost in our reporting?
Yes, as its own line if it's material. Folding a fast-growing, usage-driven cost into a general hosting line makes it much harder for anyone reading the statement to see that the two are behaving completely differently, and that's exactly the distinction a board or investor needs to see.
What's the fastest way to actually improve this margin?
Model routing, which sends each request to the cheapest model that can handle it well, usually moves gross margin faster than infrastructure optimization. The gap between a frontier model's price and a smaller model's price is often larger than any hosting efficiency gain available to you.
Sources
Where we quote a benchmark, we show its source. Other figures in this guide are estimates or general guidance, so check them against your own numbers.
- Gross margin by industry (US). NYU Stern (Aswath Damodaran), Operating and Net Margins by Industry, US, 2026.
Related Guides
What Evaluating Your AI Agent Actually Costs to Run
See where AI agent evaluation cost comes from: judge-model calls, human review and test set upkeep, with a worked run example and ways to keep spend in check.
SaaS Chart of Accounts: Numbering, COGS Split and Dimensions
Build a SaaS chart of accounts step by step: number ranges, what goes in COGS, revenue and deferred revenue accounts, and departments instead of duplicates.
Build Your Own AI Inference Cost Model in Three Tabs
How to structure a spreadsheet that turns token usage into a real cost per customer, so you can see GPU and API spend before the invoice arrives.
What Happens to Your Margin When Your Model Provider Raises Prices
Why AI wrapper products are exposed to upstream price hikes, how to see the exposure coming, and the contract and product levers that protect margin.
Tracking R&D Capitalization for AI Engineering Work That Doesn't Look Like Software
How to track engineering time and TCO for AI development work so R&D software capitalization holds up, when the work looks like training runs, not code.
Edge vs Cloud AI Inference: When On-Device Actually Pays Off
How to find your own crossover point between on-device AI inference and a cloud API, once you count hardware, model limits, and update infrastructure.