One gateway for 100+ LLMs.

Use any model through Onyx and track usage and spend in one place.

3

wire protocols

1

endpoint, one bill

0

provider keys on clients

Claude Code

Codex CLI

Cursor

Your app

Onyx Gateway

/api/gateway/v1

Metered spend

$4.2130

OpenAI

Anthropic

Gemini

Azure

Bedrock

DeepSeek

100+ models.

Zero new API keys.

Claude Fable 5
Claude Opus 4.8
Claude Sonnet 5
GPT-5.6 Sol
GPT-5.6 Terra
GPT-5.6 Luna
GPT-5.5
Gemini 3.5 Flash
Gemma 4 31B
Gemma 4 12B
Muse Spark 1.1
Grok 4.5
Mistral Small 4
Mistral Medium 3.5
Command A+
DeepSeek-V4-Pro
DeepSeek-V4-Flash
Qwen3.6-35B-A3B
Qwen3.7-Max
Qwen3.6-27B
Kimi K3
GLM-5.2
MiniMax M3
MiniMax M2.7
Step-3.7-Flash
MiMo-V2.5-Pro
Hunyuan Hy3
Nemotron 3 Ultra
Nemotron 3 Super
Nemotron 3 Nano
ERNIE 5.1
Ling-2.6-1T
Nova 2 Pro
Kimi K2.6
GLM-4.7
Qwen 3.5
DeepSeek R1
Gemma 3 27B
Llama 4 Maverick
GPT-oss 120B
Nemotron Ultra 253B
DeepSeek V3.2
Nemotron Super 49B
Step-3.5-Flash
Step3
MiMo-V2-Flash
Nemotron Nano 30B
Mistral Large 3
Llama 3.3 70B
DS-R1-Distill-Llama-70B
Llama 4 Scout
Hunyuan 2.0
Qwen3-Coder-Next
Mistral Small 3.1
GPT-oss 20B
Phi-4
Phi-4-mini
Gemma 3 12B
DS-R1-Distill-Qwen-32B
DS-R1-Distill-Qwen-14B
Gemini 3.1 Pro
Grok 3
Devstral-2-123B
Qwen3.5-35B-A3B
Qwen3.5-27B
Qwen3.5-122B-A10B
Qwen3.5-9B
Qwen3.5-4B
Ministral 14B
Ministral 8B
Ministral 3B
GLM-Z1-32B
GLM-Z1-9B
DeepSeek-R1-0528
DeepSeek-R1-0528-Qwen3-8B
Kimi-Linear-48B-A3B
Qwen3-Coder-480B-A35B

Everything a gateway should do

Three protocols. One base URL.

POST /gateway/v1/chat/completionsOpenAI Chat
POST /gateway/v1/responsesOpenAI Responses
POST /gateway/v1/messagesAnthropic Messages

Every provider you run

OpenAIAnthropicAzure OpenAIAmazon BedrockGoogle Vertex AIOpenRouterOllamavLLMLM StudioAny OpenAI-compatible endpoint

Budgets checked before the call

Token or dollar ceilings per workspace, group, or user.

Provider keys stay server-side

Clients hold a scoped Onyx token. Requests carry their permissions.

Usage and spend, in one place

Chat, agents, and gateway calls in one ledger. Search by user, drill into a day, export CSV or PDF.

Per user

daily detail

Per model

negotiated rates

Per flow

chat, agents, gateway

CSV + PDF

exports

Workspace usage

Last 30 days

Total spend

$4,218.40

Gateway share

61%

Active users

184
m.chen
$612.18
a.okafor
$498.05
s.patel
$377.90
j.laurent
$240.33
r.kim
$188.76

Gateway

Chat

Agents

Set up in minutes

01

Add models in Onyx

Access rules carry over.

02

Create a gateway token

Scoped. Revoke any time.

03

Point your tool at it

Or paste the agent prompt.

Set up the Onyx LLM Gateway for this project.
1. Ask me for my Onyx gateway base URL (https://<my-onyx-domain>/api/gateway/v1)
and a gateway-scoped personal access token (onyx_pat_...). Never echo the token.
2. Call GET <base_url>/models with "Authorization: Bearer <token>" and list the
model IDs. They look like <provider-id>/<model-name>, for example 18/claude-sonnet-5.
3. Configure the tools I use to route through the gateway:
- Claude Code: ANTHROPIC_BASE_URL, ANTHROPIC_AUTH_TOKEN, ANTHROPIC_MODEL
- Codex CLI: OPENAI_BASE_URL, OPENAI_API_KEY
- OpenAI or Anthropic SDKs in this repo: base_url and api_key in the client config
Put secrets in the environment or a gitignored .env, never in committed files.
4. Send one short test request, confirm streaming works, and tell me which model answered.

Get Started

Today

Start your 14-day free trial of Onyx (no credit card required)

Try Onyx Cloud

Secure

Enterprise-grade security and compliance. Flexible and secure deployment options.

Open-Source

Deep customizability optimized for your use case, supported by a large Open-Source community.

Trusted by

top teams

LLM Gateway | Onyx AI