Why is my deployment slower than I expected?
Last updated: September 17, 2026
The usual suspects, in the order worth checking:
Cold starts — if traffic is intermittent and the deployment scales to zero (or scales down aggressively), requests after idle periods pay a cold start. Tune
scale_down_delayor keep a minimum replica. Cold starts guideUndersized instance — a model that barely fits its GPU runs slow before it runs out of memory. Check the resources guide for sizing.
Concurrency set too high — replicas juggling too many parallel requests trade latency for throughput. Tune the concurrency target.
Engine choice — for LLMs, the engine builder (TensorRT-LLM) is typically much faster than a naive transformers loop.
The troubleshooting guide has the full diagnostic, including how to read the metrics that distinguish these cases.