What does a 429 "rate limit exceeded" error mean?
Last updated: September 17, 2026
A 429 Too Many Requests response means your request exceeded a rate limit, or arrived at a moment when no capacity was available.
For Model APIs, Baseten enforces two limits per model in your workspace:
Requests per minute (RPM) — how many API calls you can make per minute.
Tokens per minute (TPM) — how many input and output tokens you can process per minute. Cached input tokens count toward this limit at full weight, even though they're billed at a discount.
Hitting either one returns a 429. If you're seeing 429s while well under your request count, you've most likely hit the token limit instead — see Why am I getting 429s when I'm below my request limit?
What to do:
Retry with exponential backoff. Occasional 429s during traffic spikes are normal and resolve on retry.
Check your current limits in the dashboard: Model APIs → select the model — your RPM and TPM limits are shown at the top of the model page.
If 429s persist under steady traffic, your workload has outgrown your current limits — request an increase.
More detail: Inference errors