AI Model Leaderboard

Ranked on real OrcaRouter production traffic and community Battle Mode votes — not vendor-reported benchmarks.

Updated 2026-09-235 evidence sourcesNew models listed within 48 h
Measured on real traffic
100.0%
Top success rate (7d)
Google: Gemini 3.1 Flash Lite Preview
155 ms
Fastest p50 latency
Orca: OrcaVerify Text 1.0 (Free)
80.1%
Top community win rate
Qwen3.7 Max (2026-05-20)

Consolidated ranking

One blended rank per model — LMArena votes, independent benchmarks, ecosystem adoption and OrcaRouter production evidence, with fixed public weights. New models enter from external evidence on day one.

RankModelCompositeReliability$/1M blendedSpeedMomentumEvidence
#1OpenAI: GPT-5.4 Pro81.199.46% · 10000ms$165.00$60 → $270 per 1M tokens132 tokens per secondAA 66.0Artificial Analysis Intelligence Index — independent composite benchmark (0–100).Win 58.7%OrcaRouter Blind Battle win rate — our own community, blind pairwise votes.
#2DeepSeek: DeepSeek V4.1 Flash78.699.90% · 2528ms$0.38$0.15 → $0.6 per 1M tokens184 tokens per secondAA 39.5Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 13.5%OpenRouter token share — real-world ecosystem adoption.
#3OpenAI: GPT-6 Astra75.899.33% · 10000ms$30.00$10 → $50 per 1M tokens55 tokens per secondAA 52.7Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 1.4%OpenRouter token share — real-world ecosystem adoption.
#4Anthropic: Claude Opus 5.573.4100.00% · 10000ms$12.00$4 → $20 per 1M tokens1784 tokens per secondAA 57.6Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 0.0%OpenRouter token share — real-world ecosystem adoption.
#5Anthropic: Claude Opus 572.699.33% · 6220ms$15.00$5 → $25 per 1M tokens84 tokens per secondAA 50.8Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 0.9%OpenRouter token share — real-world ecosystem adoption.
#6Anthropic: Claude Fable 5.172.399.21% · 7990ms$30.00$10 → $50 per 1M tokens61 tokens per secondAA 53.4Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 0.4%OpenRouter token share — real-world ecosystem adoption.
#7Z.ai: GLM 5.372.199.86% · 3909ms$2.61$1.26 → $3.96 per 1M tokens74 tokens per secondAA 44.8Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 2.4%OpenRouter token share — real-world ecosystem adoption.
#8gpt-5.6-luna72.199.08% · 1200ms$0.90$0.2 → $1.6 per 1M tokens82 tokens per secondAA 37.3Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 6.6%OpenRouter token share — real-world ecosystem adoption.
#9Z.ai: GLM 5.3 Flash71.599.55% · 6862ms$0.16$0.075 → $0.25 per 1M tokens94 tokens per secondArena 1472LMArena community Elo — blind human votes on the text arena (CC-BY-4.0).AA 41.8Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 13.9%OpenRouter token share — real-world ecosystem adoption.
#10OpenAI: GPT-5.6 Sol71.499.14% · 10000ms$12.00$4 → $20 per 1M tokens1625 tokens per secondAA 47.0Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 1.4%OpenRouter token share — real-world ecosystem adoption.

LMArena leaderboard dataset (CC-BY-4.0) — lmarena.ai · Source: OpenRouter (openrouter.ai/rankings) · SWE-bench Verified — swebench.com · Artificial Analysis — artificialanalysis.ai

Share of the top 20 by lab
openai 20.0%anthropic 20.0%z-ai 15.0%google 10.0%Others 35.0%
40%
Human preference
LMArena · Bradley–Terry
30%
Independent benchmarks
Artificial Analysis · SWE-bench
20%
Production evidence
OrcaRouter Blind Battle win rate
10%
Ecosystem adoption
OpenRouter token share

Head-to-head

Pairwise Battle Mode win rates — how often each model beats each rival head-to-head.

Benchmarks — intelligence vs price

Independent intelligence scores plotted against input price. The dashed line is the price/intelligence Pareto frontier.

Blind Battle

Two anonymous answers, one vote. Your battles feed the 20% production-evidence weight in the consolidated rank.

⚔️ Blind Battle

Two anonymous models answer the same prompt. Read both, vote the winner, then see who wrote what — your vote feeds the ranking above.

Methodology — how this leaderboard works

Every number on this page is measured, never self-reported. This is the full method behind the consolidated rank — the same rules for every model, every day.

One blended score, fixed public weights

Each model gets a composite score from four independent evidence families: blind human preference 40%, independent benchmarks 30%, OrcaRouter production evidence 20% and ecosystem adoption 10%. The weights are fixed and public — no editorial adjustments, no paid placement, no vendor submissions.

What feeds each family

Human preference comes from LMArena’s community Elo — blind pairwise votes on the text and vision arenas (CC-BY-4.0). Independent benchmarks combine Artificial Analysis (Intelligence Index plus coding, math, GPQA Diamond and τ²-Bench) with SWE-bench Verified, the share of real GitHub issues a model resolves. Production evidence is OrcaRouter’s own Blind Battle win rates. Adoption uses OpenRouter’s published token share. Success rate, latency and throughput are shown beside each model as context — never rank inputs.

How scores become comparable

Sources score on different scales (Elo, 0–100 indices, resolved-%), so each signal is standardized against the current board and mapped to 0–100: 50 is the board average and every 15 points is one standard deviation, clamped at the extremes. A model needs at least two independent evidence families to rank at all, and the board lists the top 25 by composite.

Per-use-case boards

The text, coding, vision, math, reasoning and agentic boards re-run the same blend over segment-appropriate signals — coding swaps in the coding benchmarks and SWE-bench Verified, vision uses the LMArena vision arena — and Blind Battles only ever pair models within the same segment.

Freshness and fair play

External feeds refresh daily and first-party telemetry is live; new models enter from external evidence within 48 hours of release. No self-reported vendor numbers are used anywhere, community buzz is shown as context but never feeds a rank, and every external source is attributed under the board.

Frequently asked questions

How is the AI model leaderboard ranked?

Each model gets one blended score from four evidence families — blind human preference (LMArena), independent benchmarks (Artificial Analysis, SWE-bench), OrcaRouter production reliability, and ecosystem adoption (OpenRouter) — combined with fixed public weights of 40 / 30 / 20 / 10.

Why does a brand-new model already have a rank?

New models enter the board from external evidence within 48 hours of release. The rank hardens as community battles and production traffic accumulate.

How often is the leaderboard updated?

External feeds refresh daily and OrcaRouter telemetry is live; the board shows its last-updated date in the header.

Why is a model missing from the board?

A model needs at least two independent evidence sources to rank, and the board lists the top 25 by composite score — models with thinner evidence stay unlisted until another source picks them up.

Are different use cases ranked separately?

Yes — the pills above the consolidated table switch between text, coding, vision, math, reasoning and agentic boards, and battles only pair like with like.

When does a new model’s rank harden?

Once a model carries three evidence families including OrcaRouter’s own battle and traffic data, its rank rests on the full blend rather than external evidence alone.

Where does the data come from?

LMArena (CC-BY-4.0), Artificial Analysis, SWE-bench and OpenRouter rankings, plus OrcaRouter production telemetry — all attributed under the board.

Can vendors game this ranking?

Hard to: the weights are public, no self-reported numbers are used, and blind community battles plus live traffic anchor every score.

Can I export or embed this data?

Yes — JSON and CSV export, an embeddable per-model rank badge, and an RSS feed of rank changes are free with attribution. Links are in the footer below.

New models, rank changes and drops — follow along:X · @OrcaRouterDiscord