AI Unit Economics, FinOps & Infrastructure Cost ModelingPlaybook3 min readUpdated September 2026

Where Observability Budgets Actually Leak, and How to Contain Them

Observability tools like Datadog price on volume: how much you log, how many metrics you emit, how many traces you capture. Engineering teams tend to treat "more visibility" as an unambiguous good, which is exactly how the bill grows faster than the infrastructure it's monitoring.

The fix isn't turning observability off, it's finding the specific places volume is being generated without adding real visibility, and containing those without touching the signal engineers actually rely on.

Debug logging that never got turned back down

The single most common leak is debug-level logging turned on during an incident and never turned back down afterward. It's invisible in the product and invisible in a code review, since the log statements themselves were already there; only the verbosity setting changed. A quarterly audit of log level settings by service, compared against what's actually needed in steady state, catches this reliably.

Set an expiration on any verbosity change made during an incident, not just a mental note to revert it later. A ticket or calendar reminder tied to the specific change, due a set number of days after the incident closes, survives the on-call engineer moving to a different project far better than an intention nobody wrote down.

High cardinality metrics tags

A metric tagged with something like a user ID or request ID multiplies the number of unique time series a monitoring tool has to store, often driving cost far more than the number of metrics themselves. This tends to happen when a well-meaning engineer adds a tag for debugging one specific incident and it ships to every service using that metric going forward. Reviewing tag cardinality periodically, not just metric count, is where the real spend usually is.

Tracing sample rates left at development defaults

Default tracing configuration is often set to capture a much higher percentage of requests than production actually needs for useful visibility, especially once traffic volume is well past what the default was tuned for. Lowering the sample rate for high volume, low risk endpoints while keeping full tracing on critical paths preserves the visibility that matters while cutting the volume that doesn't. Adaptive sampling, where the tool automatically captures a higher share of slow or failed requests and a lower share of ordinary fast ones, usually gets you most of this benefit without hand-tuning a rate per endpoint.

Retention windows nobody revisited

Log and trace retention defaults to a period chosen once, early on, and rarely revisited as volume and cost both grow. Most incident investigation happens within days of an event, not months, so a long retention window on high volume data is often paying to store detail nobody will ever look at again. Shortening retention on the highest volume, lowest value data categories is usually the single largest lever available without any loss of day to day visibility.

The containment move that works without cutting visibility

Route what you keep by value, not by volume: full detail and long retention on the handful of services where an outage is expensive, reduced detail and shorter retention everywhere else. This preserves the debugging depth engineers actually need on the systems that matter most, while cutting cost on the long tail of lower stakes services that were generating the bulk of the volume without a matching amount of usefulness.

Make this an explicit, written tiering rather than an implicit habit, so a new service added six months from now gets classified deliberately instead of defaulting to whatever configuration a template happened to ship with. The list of which services sit in which tier is a short document, and it's worth keeping current as the product changes.

Audit these settings each quarter:

  • Check for debug level logging that was turned on during an incident and never turned back down.
  • Look for metric tags such as user ID or request ID that multiply unique time series, and remove the ones added for a single incident.
  • Compare tracing sample rates against what production traffic actually needs, and lower them for high volume, low risk endpoints.
  • Shorten log and trace retention on high volume data that nobody reviews once the first few days after an event have passed.
  • Keep full detail and long retention only on the few services where an outage is expensive.

Why per-host and per-GB pricing need different playbooks

Observability vendors price on different units, and the containment tactic that works for one doesn't help at all against the other. A tool billed per host or per container mostly responds to infrastructure footprint: consolidating services onto fewer, larger instances or cleaning up abandoned monitoring agents on decommissioned hosts moves that bill more than trimming log volume ever will. A tool billed per gigabyte of ingested data responds to the volume levers already covered here, log level, cardinality, sample rate and retention, and largely ignores how many hosts you're running.

Check the actual contract before assuming which lever applies. Teams sometimes spend real effort tuning log verbosity against a bill that's actually driven by host count, or chase down zombie monitoring agents on a bill that's actually driven by ingested volume, and neither effort moves the number they're aiming at.

If you're on a per-host plan, an audit of monitoring agents installed on infrastructure that's since been decommissioned or downsized is usually the fastest win, since agents rarely get uninstalled as part of a standard decommissioning checklist, and each one left running quietly keeps billing as if the host it's watching still matters.

Executive Capability Standard

What Good Looks Like

Good looks like observability spend growing roughly in line with traffic and headcount, with a clear owner for the configuration driving it.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Understand what's actually driving your observability bill: log volume, metric cardinality, trace sample rate, or retention, before assuming it's all one thing.
2. Do Manually:Audit log levels and retention settings by hand across your services on a quarterly schedule.
3. Delegate:Have a platform or infrastructure engineer own observability configuration as a standing responsibility rather than an occasional cleanup project.
4. Automate:Set budget alerts inside the observability tool itself so a volume spike triggers a notification before it shows up on next month's invoice.
5. Buy:Bring in a cloud or SaaS cost management tool if observability is one of several tools whose usage-based pricing you need visibility across.

How to Get Started

Frequently Asked Questions

Will cutting log volume make incidents harder to debug?

Cutting volume indiscriminately would, but targeting debug-level logging left on by accident, high cardinality tags, and retention on data nobody reviews doesn't touch the signal engineers rely on during an incident. The goal is removing volume that isn't adding visibility, not removing visibility itself.

Should finance or engineering own observability cost?

Engineering should own the technical changes, since only it knows what debugging needs, while finance owns visibility into the trend. Finance should ask specific questions when the bill grows faster than traffic or headcount, rather than treating the tool as fixed overhead.

How often should we audit observability configuration?

A quarterly review of log levels, tag cardinality, sample rates and retention windows catches most of the drift before it compounds into a large number. Waiting for an annual budget review to notice the trend usually means catching it well after several quarters of avoidable spend, by which point the fix is the same but the sunk cost is larger.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides