What's the difference between Model APIs, dedicated deployments, and Training?

Last updated: September 17, 2026

Baseten has the following products:

  • Model APIs — OpenAI-compatible endpoints for popular open-weight LLMs (GLM, DeepSeek, Kimi, and more), hosted and managed by Baseten. This is where most people start.

  • Dedicated deployments — your model (or one from the model library) running on GPUs dedicated to you. You control the hardware, autoscaling, and configuration, and pay per minute of compute. Use this when you need a custom model or specific performance tuning.

  • Training — Bring your own container and training code with Training Jobs, or use Loops, a Python SDK for LoRA fine-tuning and RL on a curated set of base models. Checkpoints deploy directly to inference on the same platform.

A common path: start on Model APIs, move to a dedicated deployment when you need your own model or dedicated capacity, and use Training when you want to fine-tune.

Docs: How Baseten works · Model APIs · Training