Why is my first request slower than the rest?
Last updated: September 17, 2026
A few real effects, in decreasing order of impact:
Prompt caching warms up. The first request in a session pays full prefill on the entire prompt; subsequent related requests hit the cache and start faster. Session affinity (see the prompt caching article) is what makes "subsequent" work.
Reasoning models think before they speak. Time-to-first-token includes reasoning effort — a max-reasoning request can sit "silent" for a while by design. Try a lower reasoning setting if TTFT matters more than depth.
New models are still being tuned. Performance on recently launched models typically improves over the days after launch as our engineers optimize the deployment.
If you're seeing consistently slow responses (not just first-request), tell us the model, a timestamp, and your rough input size — that's a performance report we act on.