How does prompt caching work, and how do I know if I'm getting cache hits?

Last updated: September 17, 2026

Caching is automatic — no request flags, no configuration. Prompt tokens served from the KV cache bill at a discounted rate.

Checking your hit rate: every response reports it. Look at usage.prompt_tokens_details.cached_tokens

versus usage.prompt_tokens. For an aggregate view, the usage endpoint splits cached vs. uncached input tokens per key, user, and model — hit rate is cached_input_tokens ÷ input_tokens

Improving it: cache hits depend on related requests landing on the same replica. Baseten pins requests using session IDs — Claude Code, Codex, and OpenCode are recognized automatically. For anything else, send a consistent x-session-affinity

header per conversation or task (one stable ID per task; don't reuse across unrelated tasks). Full scheme with code: cached input tokens.

One caveat worth knowing: cached tokens are cheaper, but they still count toward your tokens-per-minute limit at full weight.