Do rate limits apply per model, per API key, or per workspace?
Last updated: September 17, 2026
Rate limits are enforced per model, at the workspace level. All API keys in a workspace share the same limits — creating additional keys does not create additional capacity.
Related limits on other surfaces:
Async inference has its own limits, separate from Model API limits: 12,000 requests per minute per organization on /async_predict, and 100 requests per second on the status and cancel endpoints.
Dedicated deployments don't have account rate limits. A 429 from a dedicated deployment means capacity — every replica slot was full (CAPACITY_EXCEEDED) or the queue shed load. The fix is scaling (max replicas, concurrency target), not a rate limit request.
Details: Async rate limits · Request queuing and load shedding