Is there a size limit on request payloads?

Last updated: September 17, 2026

For async requests, /async_predict payloads are capped at 256 KiB. For sync requests, there’s also an inbound request-body limit of 100 MB.

For large inputs on any endpoint — audio, images, video, documents — the robust pattern is the same: don't inline raw bytes in the request. Upload the file to object storage (S3, GCS) and pass its URL in the payload; have your model download it inside predict. Requests stay small, you avoid base64's ~33% inflation, and the pattern works identically for sync and async.

On the output side, note that Baseten doesn't store async model outputs — results are delivered to your webhook, and large outputs should be written to your own storage from inside the model.

Docs: Async inference