Is there a timeout on requests?
Last updated: September 17, 2026
Yes - three that matter:
Synchronous requests: 1200 seconds (20 minutes) by default, measured from when the request reaches a replica. Exceed it and the request is cancelled with a
504.Parked requests get their own 1200 seconds. If your deployment is scaled to zero, an incoming request waits ("parks") for a replica for up to the same duration, and then the full predict timeout starts once it's forwarded. If parking expires before a replica is ready, you get a
429.Async requests: 1 hour per inference attempt, and they can wait in queue for up to 72 hours.
Practical guidance:
Set a client-side timeout matching your actual latency needs — the server default is far more generous than most applications want.
Work that runs longer than ~20 minutes belongs on async inference: fire the request, get a request ID immediately, receive results at your webhook.
Streaming behaves differently: the
200status is sent when the stream begins, so a timeout mid-stream shows up as a dropped connection or incomplete response rather than an error code.