polygate logopolygate

Gateway

Baseten

Hosted open models, no deployment step.

Quick start

import polygate
 
response = polygate.chat(
provider="baseten",
model="deepseek-ai/DeepSeek-V3",
messages=[{"role": "user", "content": "Say hi in one word."}],
)
 
print(response.content) # "Hi!"
print(response.usage) # Usage(prompt_tokens=..., completion_tokens=...)

Names

Any of these work as provider. They all reach the same adapter, so pick whichever reads best in your code.

baseten

Configuration

VariableHolds
BASETEN_API_KEYyour API key

Any of these may hold a comma-separated list. polygate rotates across them, parks a rate-limited one for its Retry-After, and drops one the provider rejects — so a key pool is configuration, not code.

Endpoint

https://inference.baseten.co/v1

Pass base_url to point the same wire format somewhere else — a proxy, a regional endpoint, or a gateway in front of it.

Models

Addressed vendor/model. These are examples — the gateway's own catalogue is the authority.

deepseek-ai/DeepSeek-V3zai-org/GLM-4.6

Things worth knowing

  • Baseten also serves models you deploy yourself, each on its own URL. Reach those with base_url; the default here is the shared Model APIs endpoint.

Batch

Baseten has no offline batch API. To run many requests at once, use map_chat — full price, but it works everywhere and returns in seconds. See Batching & concurrency.

Something wrong or missing here? Open an issue.