AI Unit Economics, FinOps & Infrastructure Cost ModelingPlaybook3 min readUpdated September 2026

Modeling the Power Bill Behind Your GPU Cluster

The GPU lease or purchase price is the number everyone budgets for. The power and cooling behind it is the number that keeps growing quietly afterward, and it's large enough that ignoring it in your model can make a cluster look profitable on paper while it loses money in practice.

The two things that actually drive this cost are how much power the GPUs draw and how much extra power your cooling and facility overhead adds on top of that.

Vendors Covered in this Article

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

What does PUE mean for your GPU power budget?

Power usage effectiveness is the ratio of total facility power to the power the compute itself uses; say a PUE of 1.4 means your cooling, lighting, and facility overhead adds 40 percent on top of what the GPUs draw. A colocation facility will usually quote you this number, and it's worth asking for it explicitly rather than assuming it, since the difference between a well-run facility and an average one can be a meaningful share of your total power cost.

How do you turn GPU draw into a dollar figure?

Say a GPU draws 700 watts under load and you're running 500 of them continuously, that's 350 kilowatts of compute power; at a PUE of 1.4 your facility draws 490 kilowatts total, and if your utility rate is 10 cents per kilowatt-hour, that's roughly $1,176 an hour, or a bit over $850,000 a month, before you've counted a single dollar of the GPU lease itself.

To turn GPU draw into a monthly power bill, work through these steps:

  1. Take each GPU's draw under load and multiply by the number of GPUs running continuously to get total compute power.
  2. Multiply by the facility's PUE, which you should ask the colocation provider to state explicitly, to get total facility power.
  3. Multiply by your utility rate per kilowatt-hour and by the hours in the billing period to get the energy charge.
  4. Add demand charges based on your peak draw, since they can dominate the bill for spiky training workloads.
  5. Replace nameplate assumptions with real utilization from your scheduler and rebuild the model as that utilization changes.

Where utility contracts change the number

Demand charges, a separate fee based on your peak power draw in a billing period rather than total energy used, can dominate the bill for a workload that spikes rather than running flat, which most training workloads do. Negotiating a flatter draw profile, or a utility contract structured around your actual usage pattern rather than a generic industrial rate, is often worth more than shopping for a marginally cheaper per-kilowatt-hour rate.

Some utilities also offer interruptible or curtailable-load rates, a lower price in exchange for agreeing to reduce draw during grid stress events, which can fit a training workload that can tolerate being paused far better than it fits a workload serving live customer traffic. Whether that tradeoff makes sense depends entirely on how tolerant your specific workload is to an unscheduled pause, which is a technical question worth asking before you assume the discount is free money.

Building the model around actual utilization, not nameplate capacity

A model that assumes every GPU runs at full draw around the clock overstates the bill for a cluster that actually sits idle between training runs, and understates it for one running dense, back-to-back jobs. Pull your real utilization rate from your scheduler, not an assumed number, and apply it to the power model, since the gap between nameplate power draw and actual metered draw is often the single largest source of error in a first-pass estimate.

For example, imagine two clusters of identical size. One runs dense, back-to-back training jobs, while the other sits idle between occasional runs. A model built on nameplate draw gives both the same bill, which overstates the second and understates the first. A practical decision rule: build the model from metered draw for at least one full billing period, split the result into energy charges and demand charges, and flag any month where the demand portion looks unusually large. That flag tells you whether your cost problem is how much power you use or when you use it, and those two problems have different fixes.

The contract paperwork this actually depends on

A colocation or utility agreement with demand-charge structure, PUE commitments, and any power-cost pass-through clauses needs to be reviewed and signed with the same care as any other multi-year commitment this size, not rushed because it's an infrastructure decision rather than a real estate one. Foxit eSign is a reasonable place to manage the signature and amendment trail on agreements like this; Process Street can hold the actual due-diligence checklist your team runs before signing one, so the review doesn't depend on whoever did it last time remembering every clause to check.

Revisiting the model as utilization changes

A power and cooling model built at launch, when utilization is low, understates the real cost once the cluster is running closer to capacity, since fixed facility overhead gets spread across more actual compute hours. Rebuild the model against actual metered draw at least twice in the first year, not just once at the planning stage, since the assumptions that looked reasonable before launch rarely survive contact with real utilization.

Executive Capability Standard

What Good Looks Like

The standard is a power and cooling cost model built from your actual metered draw and your actual utility contract terms, rebuilt as utilization changes, not a one-time estimate from the planning stage.

Building The Capability (5-Stage Skill Ladder)

1. Learn:Ask your colocation provider or facilities team for your actual measured PUE and last three months of metered power draw.
2. Do Manually:Build a simple spreadsheet converting GPU count and utilization into a dollar figure using your real PUE and utility rate.
3. Delegate:Ask your infrastructure lead to report actual versus modeled power cost each quarter as utilization changes.
4. Automate:Pull metered power data automatically into your cost model so it updates without a manual re-entry each time.
5. Buy:Bring in an energy or facilities consultant to review your utility contract structure before signing or renewing a multi-year colocation agreement.

How to Get Started

Disclosure: We may earn a commission if you buy through some links on this page. It doesn't change what we recommend.

Frequently Asked Questions

How much does cooling actually add to a GPU cluster's power bill?

It depends heavily on the facility, but say a facility's PUE lands somewhere between 1.2 and 1.5, that means cooling and overhead are adding 20 to 50 percent on top of the GPUs' own power draw. Ask your colocation provider for their actual measured PUE rather than a marketing figure.

What's a demand charge, and why does it matter more than the per-kilowatt-hour rate?

A demand charge is a fee based on your highest power draw in a billing period, not your total usage. A workload that spikes to peak draw for even a short window can set a charge that applies to the whole period, so a flatter, more predictable draw profile can matter more than the headline electricity rate.

Should this modeling live with finance or with the infrastructure team?

Both, working from the same numbers. Infrastructure knows the actual power draw and PUE; finance knows the utility contract structure and how to model demand charges against usage patterns. Building the model in one place without the other tends to miss either the technical reality or the contract mechanics.

About the numbers

This guide doesn't quote a sourced benchmark. Figures in it are estimates or general guidance, so check them against your own numbers.

Related Guides