What does a 429 "rate limit exceeded" error mean?

Last updated: September 17, 2026

A 429 Too Many Requests response means your request exceeded a rate limit, or arrived at a moment when no capacity was available.

For Model APIs, Baseten enforces two limits per model in your workspace:

  • Requests per minute (RPM) — how many API calls you can make per minute.

  • Tokens per minute (TPM) — how many input and output tokens you can process per minute. Cached input tokens count toward this limit at full weight, even though they're billed at a discount.

Hitting either one returns a 429. If you're seeing 429s while well under your request count, you've most likely hit the token limit instead — see Why am I getting 429s when I'm below my request limit?

What to do:

Retry with exponential backoff. Occasional 429s during traffic spikes are normal and resolve on retry.

Check your current limits in the dashboard: Model APIs → select the model — your RPM and TPM limits are shown at the top of the model page.

If 429s persist under steady traffic, your workload has outgrown your current limits — request an increase.

More detail: Inference errors