What's the difference between requests per minute and tokens per minute?

Last updated: September 17, 2026

They're two separate limits, and either one can trigger a 429:

RPM (requests per minute) counts API calls. A request with a 10-token prompt and a request with a 100,000-token prompt each count as one request.

TPM (tokens per minute) counts the tokens those calls process — input and output combined. Cached and uncached input tokens count equally toward TPM, even though cached tokens cost less.

Most workloads that hit rate limits unexpectedly are TPM-bound, not RPM-bound. This is especially true for coding agents and long-context applications: a coding agent sending 20 requests per minute with 50,000-token contexts is using 1M+ tokens per minute while barely registering on the request counter.

How to tell which limit you're hitting: compare your usage against both limits. Your per-model limits are shown in the dashboard under Model APIs, and you can pull exact token and request counts (in 1-minute buckets, split by cached/uncached input and output) from the usage endpoint.

Details: Pricing and limits