Gateway
DeepInfra
Open models billed per token.
Quick start
import polygate response = polygate.chat( provider="deepinfra", model="meta-llama/Llama-3.3-70B-Instruct", messages=[{"role": "user", "content": "Say hi in one word."}],) print(response.content) # "Hi!"print(response.usage) # Usage(prompt_tokens=..., completion_tokens=...)Names
Any of these work as provider. They all reach the same adapter, so pick whichever reads best in your code.
deepinfraConfiguration
| Variable | Holds |
|---|---|
| DEEPINFRA_API_KEY | your API key |
Any of these may hold a comma-separated list. polygate rotates across them, parks a rate-limited one for its Retry-After, and drops one the provider rejects — so a key pool is configuration, not code.
Endpoint
https://api.deepinfra.com/v1/openai
Pass base_url to point the same wire format somewhere else — a proxy, a regional endpoint, or a gateway in front of it.
Models
Addressed vendor/model. These are examples — the gateway's own catalogue is the authority.
meta-llama/Llama-3.3-70B-InstructThings worth knowing
- The OpenAI-compatible surface lives under
/v1/openai, beside DeepInfra's own native API at/v1/inference.
Batch
DeepInfra has no offline batch API. To run many requests at once, use map_chat — full price, but it works everywhere and returns in seconds. See Batching & concurrency.