Ranked on real OrcaRouter production traffic and community Battle Mode votes — not vendor-reported benchmarks.
Blind Battle Mode votes, ranked by a Bradley–Terry (Elo) model. Dots show the rating; bars show the 95% confidence interval. Models within a band are statistically tied.
Two anonymous models answer the same prompt. Read both, vote the winner, then see who wrote what — your vote feeds the ranking above.
| Rank | Model | Arena Rating | Method | Votes | W·L·T |
|---|---|---|---|---|---|
| #1 | Qwen3.7 Max (2026-05-20) | blind · BT | 168 | 129·32·7 | |
| #2 | DEDeepSeek: DeepSeek V4 Pro | blind · BT | 250 | 161·77·12 | |
| OpenAI: GPT-5.4 Pro | blind · BT | 281 | 158·111·12 | ||
| OpenAI: GPT-5.5 | blind · BT | 230 | 109·109·12 | ||
| Qwen: Qwen3.5-35B-A3B | blind · BT | 310 | 149·149·12 | ||
| Qwen3.7 Max | blind · BT | 224 | 84·130·10 | ||
| Anthropic: Claude Opus 4.7 | blind · BT | 247 | 117·117·13 | ||
| Qwen: Qwen3.6 35B A3B | blind · BT | 260 | 124·124·12 | ||
| Anthropic: Claude Opus 4.8 | blind · BT | 245 | 116·116·13 | ||
| Google: Gemma 4 26B A4B | blind · BT | 242 | 114·116·12 | ||
| Qwen: Qwen3.6 Plus | blind · BT | 278 | 108·158·12 | ||
| #12 | KIMoonshotAI: Kimi K2.7 Code | blind · BT | 224 | 119·93·12 | |
| MIMiniMax: MiniMax M3 | blind · BT | 265 | 127·126·12 | ||
| Google: Gemma 4 31B | blind · BT | 273 | 112·149·12 | ||
| OpenAI: GPT-5.5 Pro | blind · BT | 294 | 95·187·12 | ||
| Qwen: Qwen3.5-27B | blind · BT | 213 | 100·101·12 | ||
| ZAZ.ai: GLM 5.2 | blind · BT | 235 | 141·82·12 | ||
| #18 | MIMiniMax: MiniMax M2.7 | blind · BT | 226 | 81·133·12 | |
| Google: Nano Banana 2 (Gemini 3.1 Flash Image Preview) | blind · BT | 281 | 151·118·12 | ||
| OpenAI: GPT-5.4 | blind · BT | 245 | 117·116·12 | ||
| Google: Gemini 3.1 Flash Lite Preview | blind · BT | 285 | 137·136·12 | ||
| MIMiniMax M2.7 highspeed | blind · BT | 246 | 117·117·12 | ||
| OpenAI: GPT-5.4 Nano | blind · BT | 308 | 196·100·12 | ||
| Qwen: Qwen3.7 Plus | blind · BT | 248 | 118·118·12 | ||
| ZAZ.ai: GLM 5.1 | blind · BT | 220 | 104·104·12 | ||
| KLKling: Kling 3.0 Turbo | blind · BT | 256 | 121·123·12 | ||
| Qwen: Qwen3.6 Flash | blind · BT | 188 | 71·105·12 | ||
| #28 | Gemini 3.5 Flash | blind · BT | 285 | 96·177·12 | |
| OpenAI: GPT-5.4 Mini | blind · BT | 206 | 75·122·9 | ||
| DEDeepSeek: DeepSeek V4 Flash | blind · BT | 128 | 61·61·6 |
Pairwise Battle Mode win rates — how often each model beats each rival head-to-head.
The strongest model in each capability — coding, math, reasoning and more — with a per-axis profile against the field median.
Independent intelligence scores plotted against input price. The dashed line is the price/intelligence Pareto frontier.
Measured on OrcaRouter production traffic over the last 7 days.
| Model | Success % (7d) | p50 | p99 | Err % | $/1M (in→out) | tok/s |
|---|---|---|---|---|---|---|
| openai/gpt-4.1-2025-04-14 | 100.0% | 3.67 s | 7.44 s | 0.0% | $2.00 → $8.00 | 87 |
| Qwen: Qwen3 VL 235B A22B Instruct | 100.0% | 10.00 s | 10.00 s | 0.0% | $0.40 → $1.60 | 74 |
| openai/gpt-3.5-turbo-0125 | 100.0% | 10.00 s | 10.00 s | 0.0% | $0.50 → $1.50 | 1512 |
| openai/gpt-3.5-turbo-1106 | 100.0% | 1.67 s | 4.15 s | 0.0% | $1.00 → $2.00 | 145 |
| OpenAI: GPT-4 | 100.0% | 6.88 s | 8.12 s | 0.0% | $30.00 → $60.00 | 294 |
| OpenAI: GPT-5 | 100.0% | 10.00 s | 10.00 s | 0.0% | $1.25 → $10.00 | 1504 |
| openai/gpt-5.1-chat-latest | 100.0% | 2.38 s | 3.50 s | 0.0% | $1.25 → $10.00 | 110 |
| MIMiniMax M2.7 highspeed | 100.0% | 3.67 s | 3.67 s | 0.0% | $0.60 → $2.40 | 83 |
| OpenAI: GPT-4o (2024-11-20) | 100.0% | 1.83 s | 4.34 s | 0.0% | $2.50 → $10.00 | 177 |
| OpenAI: GPT-5.1-Codex | 100.0% | 1.00 s | 1.83 s | 0.0% | $1.25 → $10.00 | 87 |
| Qwen: Qwen3.5-122B-A10B | 100.0% | 5.00 s | 10.00 s | 0.0% | $0.12 → $0.92 | 89 |
| Qwen: Qwen3.6 Flash | 100.0% | 3.85 s | 10.00 s | 0.0% | $0.25 → $1.50 | 303 |
| qwen/qwen3.6-plus-2026-04-02 | 100.0% | 4.18 s | 7.86 s | 0.0% | $0.28 → $1.65 | 56 |
| MIMiniMax: MiniMax M2.7 | 100.0% | 1.32 s | 10.00 s | 0.0% | $0.30 → $1.20 | 81 |
| OpenAI: GPT-4.1 Nano | 100.0% | 1.36 s | 5.00 s | 0.0% | $0.10 → $0.40 | 99 |
| OpenAI: GPT-4o | 100.0% | 902 ms | 10.00 s | 0.0% | $2.50 → $10.00 | 98 |
| Qwen3.7 Max | 100.0% | 3.89 s | 10.00 s | 0.0% | $1.25 → $3.75 | 56 |
| OpenAI: GPT-5.3-Codex | 100.0% | 1.10 s | 2.89 s | 0.0% | $1.75 → $14.00 | 62 |
| Qwen: Qwen3.7 Plus | 100.0% | 3.96 s | 10.00 s | 0.0% | $0.35 → $1.42 | 58 |
| MIMiniMax: MiniMax M2.5 | 100.0% | 3.50 s | 4.63 s | 0.0% | $0.30 → $1.20 | 100 |
| openai/gpt-5.2-2025-12-11 | 100.0% | 2.50 s | 10.00 s | 0.0% | $1.75 → $14.00 | 74 |
| KIkimi/kimi-k2.6 | 100.0% | 3.92 s | 10.00 s | 0.0% | $0.95 → $4.00 | 36 |
| OpenAI: GPT-3.5 Turbo 16k | 100.0% | 5.00 s | 7.07 s | 0.0% | $3.00 → $4.00 | 103 |
| OpenAI: GPT-4o-mini (2024-07-18) | 100.0% | 5.00 s | 10.00 s | 0.0% | $0.15 → $0.60 | 32 |
| openai/gpt-4o-mini-search-preview-2025-03-11 | 100.0% | 5.19 s | 5.19 s | 0.0% | $0.15 → $0.60 | — |
Which models actually carry OrcaRouter production traffic, ranked by token throughput over the selected window.
We have seen — models so far this period. The volume ranking unlocks once enough traffic has accumulated to rank models fairly.
In the meantime, the Overall and Reliability boards above are already live.
How closely the community ranking tracks independent external rankings, measured by Spearman’s ρ and Kendall’s τ rank correlation.
Each model’s Arena Rating over the last 90 days, from daily snapshots. A regression badge flags a sustained drop.
What the community is talking about — mention volume, 14-day momentum and sentiment. Context only, never a ranking signal.