Why did my deployment fail to get capacity?

Last updated: September 17, 2026

Two different situations that look similar:

At deploy time — the deployment can't provision hardware (stays in "requesting"). That's pool availability; see My deployment is stuck at "requesting."

At request time — a running deployment returns 429 with CAPACITY_EXCEEDED, or sometimes 529depending on the backpressure threshold. That means the deployment couldn’t accept more work at that moment — for example, replicas were saturated, the queue was full, or the deployment was still scaling up.

This is usually fixable in your configuration:

  • Raise max replicas so the autoscaler can add capacity under load

  • Tune the concurrency target so each replica accepts the right amount of parallel work

  • Make sure your client retries with backoff — brief capacity spikes are normal during scale-up

Docs: Request queuing and load shedding