How should I load test my deployment before production?
Last updated: September 17, 2026
A load test on Baseten is really testing your autoscaling configuration, not just the model. A sequence that works:
Benchmark one replica first. Send representative requests (real payload sizes, real prompt lengths) at increasing concurrency to a deployment pinned at one replica, and find where latency degrades. That measurement is how you set the concurrency target — at measured capacity, not above it.
Then test scaling behavior. Raise max replicas, ramp traffic the way production will (gradual vs. spike), and watch queue time and replica count in the metrics dashboard. Requests queuing while replicas sit at the ceiling → raise max replicas. Replicas scaling but latency still poor → revisit the concurrency target.
Account for cold starts. A test that starts from zero replicas measures cold start + inference, not steady state. Test both, deliberately — production will experience both.
Pre-warm capacity for the test itself. One-time autoscaling schedules exist for exactly this — setting a replica floor for a planned load test or launch window, so you measure inference rather than provisioning.
For very large tests (sustained high volume, big GPU counts), give us a heads-up first — we can make sure shared pool capacity doesn't skew your numbers.
Docs: Autoscaling · Traffic patterns