What does running an AI application really cost?

Build an AI application cost model that includes compute, idle capacity, supporting services and operational work.

Electrical distribution panels and a power meter at the NERSC data centre.

EDITORIAL DRAFT — outline only. Not ready for publication. AI infrastructure decisions, part 4 of 6. Audience: teams investigating AI application spending.

The compute bill is only part of the cost

[Open with a clearly labelled example. Explain why a token rate or GPU hourly price cannot by itself predict the application’s monthly cost.]

Build the full cost model

  • [Inference/training compute, idle capacity, retries and peak demand.]
  • [Storage, network transfer, retrieval, logging and monitoring.]
  • [Engineering, support, security, evaluation and incident handling.]
  • [Commitments, minimum charges, taxes and one-off migration costs.]

Compare scenarios, not one precise-looking number

[Show low, expected and peak usage scenarios. State the unit of useful work, quality target, latency constraints and what changes when traffic or context length grows. Identify the assumptions that dominate the result.]

Evidence and assumptions to publish

[Record hardware, model, precision, utilisation, workload and measurement date. Source prices with region, currency and retrieval date. Separate measured usage from estimates, and list exclusions. Use synthetic or permissioned, anonymised examples—never confidential customer bills.]

Reader next step

[Invite readers to request a preliminary cost discussion through the AI hosting requirements review on /services/. Ask for approximate spend, workload and the problem—not credentials or sensitive billing records. Explain that scope and any paid assessment fee are agreed first; do not promise savings or an automated cost report.]