Gateway
Nebius AI Studio
Open models from a European cloud.
Quick start
import polygate response = polygate.chat( provider="nebius", model="meta-llama/Llama-3.3-70B-Instruct", messages=[{"role": "user", "content": "Say hi in one word."}],) print(response.content) # "Hi!"print(response.usage) # Usage(prompt_tokens=..., completion_tokens=...)Names
Any of these work as provider. They all reach the same adapter, so pick whichever reads best in your code.
nebiusConfiguration
| Variable | Holds |
|---|---|
| NEBIUS_API_KEY | your API key |
Any of these may hold a comma-separated list. polygate rotates across them, parks a rate-limited one for its Retry-After, and drops one the provider rejects — so a key pool is configuration, not code.
Endpoint
https://api.studio.nebius.com/v1
Pass base_url to point the same wire format somewhere else — a proxy, a regional endpoint, or a gateway in front of it.
Models
Addressed vendor/model. These are examples — the gateway's own catalogue is the authority.
meta-llama/Llama-3.3-70B-InstructThings worth knowing
- The host is
api.studio.nebius.com, notapi.nebius.com— the latter is the cloud control plane and will not serve inference.
Batch
Nebius AI Studio has no offline batch API. To run many requests at once, use map_chat — full price, but it works everywhere and returns in seconds. See Batching & concurrency.