What does it cost to run a model?
Last updated: September 17, 2026
Depends on the product:
Model APIs bill per million tokens, with per-model rates on the pricing page. Cached input tokens are billed at a discounted rate automatically.
Dedicated deployments bill per minute of compute, by instance type. Rates for every GPU/CPU configuration are in the instance type reference. With scale-to-zero enabled, you don’t pay for compute while the deployment has no active replicas.
Training bills per minute of GPU time while the job runs.
Usage is metered hourly and visible anytime in your billing dashboard.