One gateway. Every model.

Secure zero-markup inference on pay-as-you-go and subscriptions. Or bring your own keys. Lower your bill with adaptive routing.

No credit card · live in 60 seconds

200+
models, one endpoint
0%
token markup, ever
75.5%
routing accuracy
beats GPT-5 on RouterArena
<50ms
mid-stream failover
40+%
lower inference cost
with adaptive session-aware routing
<1%
quality degradation
with automatic frontier escalation
All integrations
from openai import OpenAI
 
client = OpenAI(
base_url="https://api.orcarouter.ai/v1",
api_key=ORCAROUTER_API_KEY,
)
resp = client.chat.completions.create(
model="orcarouter/auto", # we grade + route
messages=[{"role": "user", "content": "..."}],
)

One line. We grade each prompt, route to frontier or OSS, and add $0.

The gateway

Route smarter. Ship safer. Spend less.

Route

Every prompt graded, then sent to the model that answers it best.

Browse models

Observe

Full logs, live spend and a per-prompt receipt. No black boxes.

See a receipt

Manage

Version prompts and reuse cached calls without touching code.

Version a prompt

Govern

Guardrails and an agent firewall that stop things, not just log them.

See the firewall
Setup

Live in 60 seconds.

Step 1

Point your SDK at us

Set base_url to api.orcarouter.ai/v1 and swap your API key. No other code changes needed.

Step 2

We route, guard and observe

Graded in under 1ms, with failover, caching and full logs built in.

Step 3

You ship, on one endpoint

Direct to each provider at their published rate — we add $0 per token.

Models

Every model. One price list.

Live, side-by-side pricing — what you'd pay the provider directly.

View all 200+ →
ModelRouted toInput /MOutput /MContextQuality
qwen/qwen3.8-27b-freeNEWAlibaba Cloud262K4.0
qwen/qwen3.8-27bNEWAlibaba Cloud$0.330$2.40262K4.0
deepseek/deepseek-v4-pro-0813NEWDeepSeek$0.442$0.8841M7.0
grok/grok-4.6NEW$2.00$6.00500K9.0
meta/muse-spark-1.2NEW$1.25$4.251M8.0
qwen/qwen3.8-maxNEWAlibaba Cloud$2.00$6.001M9.0
deepseek/deepseek-v4-flash-0731NEWDeepSeek$0.147$0.2951M6.0
minimax/minimax-h3NEW$0.080 / second5.0
+ 194 more models · prices update every 60 seconds
Paying for tokens

Three ways to pay for tokens.

Top-ups and subscriptions bill at provider price. We add $0.

Plans below are for team features. They don't change token price.

Pricing

We never take a cut of your token spend.

Our revenue comes from optional team features.

Hacker

Free
Forever. Zero markup on all tokens.
Route — 200+ models, auto-failover
Observe — basic dashboard
Manage — prompt versioning
3 API keys · 0% token markup
Start free

Enterprise

Custom
SLA commitments and private deployment.
Everything in Team
Private / on-prem deployment
99.99% uptime SLA
Dedicated infrastructure
Dedicated support & custom pricing
Trust & Compliance
Independently audited, continuously compliant — reports available under NDA.
From the blog

Fresh from the engine room.

What we're building and why — our latest posts.

All posts →
FAQ

AI routers, LLM gateways, adaptive routing — answered.

What is an AI router?

An AI router sits between your app and many models and picks the best model for each request. OrcaRouter grades every prompt in under 1 ms and routes it — frontier for hard reasoning, open-source for routine — across 200+ models at zero token markup.

What is an LLM router and how does it work?

An LLM router classifies each prompt and matches it to the model most likely to answer well at the lowest cost. OrcaRouter routes on contextual embeddings with online learning from live traffic — 75.5% accuracy on the public RouterArena leaderboard.

Is OrcaRouter an AI gateway?

Yes. OrcaRouter is a production AI gateway: one OpenAI-compatible endpoint with adaptive routing, load balancing, automatic failover, guardrails, an agent firewall, prompt caching and per-request observability.

What is an AI CDN?

Like a CDN serves content from the best edge, an AI CDN serves inference from the best provider: healthy, fast, cheap capacity, cached repeated prompts, failover on outages. OrcaRouter plays this role across 200+ models.

What is adaptive routing?

Adaptive routing learns the quality/cost trade-off from your own traffic. Point each workspace at cheapest-that-clears-the-bar, highest quality, or balanced — or let orcarouter/auto keep tuning the choice per request.

Do I need both an AI gateway and an LLM router?

They solve different problems — governance vs model choice — but you don't need two systems: OrcaRouter routes per request and enforces budgets, guardrails and observability on the same hop.

Does routing add latency?

Prompt grading takes under 1 ms and total added latency stays under 50 ms — usually won back many times over by faster providers and cached prompt tokens.

Does OrcaRouter mark up token prices?

No. You pay each provider's published rate; OrcaRouter adds $0 per token and monetizes optional Team and Enterprise features.

Can I keep my existing OpenAI SDK code?

Yes — switch base_url to https://api.orcarouter.ai/v1 and your OpenAI, Anthropic or Google SDK code keeps working: Chat Completions, Responses, Embeddings, Images, Audio and streaming.

What happens when a provider goes down?

OrcaRouter retries against healthy fallback capacity for the same or an equivalent model before the response starts, so upstream outages don't surface to your users.

Smarter, safer, cost-efficient.

Swap one line. That's the migration.

Beats GPT-5 & Azure on RouterArenaBacked by published research

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube