AI Model Leaderboard

Ranked on real OrcaRouter production traffic and community Battle Mode votes — not vendor-reported benchmarks.

Measured on real traffic
100.0%
Top success rate (7d)
openai/gpt-4.1-2025-04-14
399 ms
Fastest p50 latency
DeepSeek: DeepSeek V3
80.1%
Top community win rate
Qwen3.7 Max (2026-05-20)

Overall — community win rate

Blind Battle Mode votes, ranked by a Bradley–Terry (Elo) model. Dots show the rating; bars show the 95% confidence interval. Models within a band are statistically tied.

Style control
Coming soon — needs more battle data

⚔️ Blind Battle

Two anonymous models answer the same prompt. Read both, vote the winner, then see who wrote what — your vote feeds the ranking above.

RankModelArena RatingMethodVotesW·L·T
#1Qwen3.7 Max (2026-05-20)1746±140
blind · BT
168129·32·7
#2DeepSeek: DeepSeek V4 Pro1526±76
blind · BT
250161·77·12
OpenAI: GPT-5.4 Pro1523±81
blind · BT
281158·111·12
OpenAI: GPT-5.51521±81
blind · BT
230109·109·12
Qwen: Qwen3.5-35B-A3B1519±88
blind · BT
310149·149·12
Qwen3.7 Max1519±125
blind · BT
22484·130·10
Anthropic: Claude Opus 4.71518±103
blind · BT
247117·117·13
Qwen: Qwen3.6 35B A3B1518±91
blind · BT
260124·124·12
Anthropic: Claude Opus 4.81516±96
blind · BT
245116·116·13
Google: Gemma 4 26B A4B1516±91
blind · BT
242114·116·12
Qwen: Qwen3.6 Plus1515±125
blind · BT
278108·158·12
#12MoonshotAI: Kimi K2.7 Code1301±69
blind · BT
224119·93·12
MiniMax: MiniMax M31300±59
blind · BT
265127·126·12
Google: Gemma 4 31B1300±58
blind · BT
273112·149·12
OpenAI: GPT-5.5 Pro1299±68
blind · BT
29495·187·12
Qwen: Qwen3.5-27B1298±63
blind · BT
213100·101·12
Z.ai: GLM 5.21298±67
blind · BT
235141·82·12
#18MiniMax: MiniMax M2.71081±66
blind · BT
22681·133·12
Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview)1079±105
blind · BT
281151·118·12
OpenAI: GPT-5.41078±73
blind · BT
245117·116·12
Google: Gemini 3.1 Flash Lite Preview1078±98
blind · BT
285137·136·12
MiniMax M2.7 highspeed1077±106
blind · BT
246117·117·12
OpenAI: GPT-5.4 Nano1077±102
blind · BT
308196·100·12
Qwen: Qwen3.7 Plus1076±90
blind · BT
248118·118·12
Z.ai: GLM 5.11076±78
blind · BT
220104·104·12
Kling: Kling 3.0 Turbo1074±91
blind · BT
256121·123·12
Qwen: Qwen3.6 Flash1073±70
blind · BT
18871·105·12
#28Gemini 3.5 Flash852±123
blind · BT
28596·177·12
OpenAI: GPT-5.4 Mini851±123
blind · BT
20675·122·9
DeepSeek: DeepSeek V4 Flash850±128
blind · BT
12861·61·6

Head-to-head

Pairwise Battle Mode win rates — how often each model beats each rival head-to-head.

Category leaders

The strongest model in each capability — coding, math, reasoning and more — with a per-axis profile against the field median.

Benchmarks — intelligence vs price

Independent intelligence scores plotted against input price. The dashed line is the price/intelligence Pareto frontier.

Reliability & cost

Measured on OrcaRouter production traffic over the last 7 days.

Model
Success % (7d)
p50
p99
Err %
$/1M (in→out)
tok/s
openai/gpt-4.1-2025-04-14100.0%3.67 s7.44 s
0.0%
$2.00 → $8.0087
Qwen: Qwen3 VL 235B A22B Instruct100.0%10.00 s10.00 s
0.0%
$0.40 → $1.6074
openai/gpt-3.5-turbo-0125100.0%10.00 s10.00 s
0.0%
$0.50 → $1.501512
openai/gpt-3.5-turbo-1106100.0%1.67 s4.15 s
0.0%
$1.00 → $2.00145
OpenAI: GPT-4100.0%6.88 s8.12 s
0.0%
$30.00 → $60.00294
OpenAI: GPT-5100.0%10.00 s10.00 s
0.0%
$1.25 → $10.001504
openai/gpt-5.1-chat-latest100.0%2.38 s3.50 s
0.0%
$1.25 → $10.00110
MiniMax M2.7 highspeed100.0%3.67 s3.67 s
0.0%
$0.60 → $2.4083
OpenAI: GPT-4o (2024-11-20)100.0%1.83 s4.34 s
0.0%
$2.50 → $10.00177
OpenAI: GPT-5.1-Codex100.0%1.00 s1.83 s
0.0%
$1.25 → $10.0087
Qwen: Qwen3.5-122B-A10B100.0%5.00 s10.00 s
0.0%
$0.12 → $0.9289
Qwen: Qwen3.6 Flash100.0%3.85 s10.00 s
0.0%
$0.25 → $1.50303
qwen/qwen3.6-plus-2026-04-02100.0%4.18 s7.86 s
0.0%
$0.28 → $1.6556
MiniMax: MiniMax M2.7100.0%1.32 s10.00 s
0.0%
$0.30 → $1.2081
OpenAI: GPT-4.1 Nano100.0%1.36 s5.00 s
0.0%
$0.10 → $0.4099
OpenAI: GPT-4o100.0%902 ms10.00 s
0.0%
$2.50 → $10.0098
Qwen3.7 Max100.0%3.89 s10.00 s
0.0%
$1.25 → $3.7556
OpenAI: GPT-5.3-Codex100.0%1.10 s2.89 s
0.0%
$1.75 → $14.0062
Qwen: Qwen3.7 Plus100.0%3.96 s10.00 s
0.0%
$0.35 → $1.4258
MiniMax: MiniMax M2.5100.0%3.50 s4.63 s
0.0%
$0.30 → $1.20100
openai/gpt-5.2-2025-12-11100.0%2.50 s10.00 s
0.0%
$1.75 → $14.0074
kimi/kimi-k2.6100.0%3.92 s10.00 s
0.0%
$0.95 → $4.0036
OpenAI: GPT-3.5 Turbo 16k100.0%5.00 s7.07 s
0.0%
$3.00 → $4.00103
OpenAI: GPT-4o-mini (2024-07-18)100.0%5.00 s10.00 s
0.0%
$0.15 → $0.6032
openai/gpt-4o-mini-search-preview-2025-03-11100.0%5.19 s5.19 s
0.0%
$0.15 → $0.60
Showing 1 to 25 of 133

Volume — production traffic

Which models actually carry OrcaRouter production traffic, ranked by token throughput over the selected window.

Volume ranking is warming up

We have seen — models so far this period. The volume ranking unlocks once enough traffic has accumulated to rank models fairly.

Models seen

In the meantime, the Overall and Reliability boards above are already live.

Validation — agreement with external rankings

How closely the community ranking tracks independent external rankings, measured by Spearman’s ρ and Kendall’s τ rank correlation.

Buzz — popularity & sentiment

What the community is talking about — mention volume, 14-day momentum and sentiment. Context only, never a ranking signal.