Ranked on real OrcaRouter production traffic and community Battle Mode votes — not vendor-reported benchmarks.
Updated 2026-08-095 evidence sourcesNew models listed within 48 h
Measured on real traffic
100.0%
Top success rate (7d)
OpenAI: GPT-3.5 Turbo
450 ms
Fastest p50 latency
OpenAI: GPT-5.1-Codex-Mini
80.1%
Top community win rate
Qwen3.7 Max (2026-05-20)
Consolidated ranking
One blended rank per model — LMArena votes, independent benchmarks, ecosystem adoption and OrcaRouter production evidence, with fixed public weights. New models enter from external evidence on day one.
Arena 1494LMArena community Elo — blind human votes on the text arena (CC-BY-4.0).AA 62.1Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 0.3%OpenRouter token share — real-world ecosystem adoption.
Arena 1480LMArena community Elo — blind human votes on the text arena (CC-BY-4.0).AA 51.6Artificial Analysis Intelligence Index — independent composite benchmark (0–100).OR 3.5%OpenRouter token share — real-world ecosystem adoption.
Pairwise Battle Mode win rates — how often each model beats each rival head-to-head.
Benchmarks — intelligence vs price
Independent intelligence scores plotted against input price. The dashed line is the price/intelligence Pareto frontier.
Blind Battle
Two anonymous answers, one vote. Your battles feed the 20% production-evidence weight in the consolidated rank.
⚔️ Blind Battle
Two anonymous models answer the same prompt. Read both, vote the winner, then see who wrote what — your vote feeds the ranking above.
Methodology — how this leaderboard works
Every number on this page is measured, never self-reported. This is the full method behind the consolidated rank — the same rules for every model, every day.
One blended score, fixed public weights
Each model gets a composite score from four independent evidence families: blind human preference 40%, independent benchmarks 30%, OrcaRouter production evidence 20% and ecosystem adoption 10%. The weights are fixed and public — no editorial adjustments, no paid placement, no vendor submissions.
What feeds each family
Human preference comes from LMArena’s community Elo — blind pairwise votes on the text and vision arenas (CC-BY-4.0). Independent benchmarks combine Artificial Analysis (Intelligence Index plus coding, math, GPQA Diamond and τ²-Bench) with SWE-bench Verified, the share of real GitHub issues a model resolves. Production evidence is OrcaRouter’s own: Blind Battle win rates plus live gateway telemetry — success rate, latency and routed-token volume. Adoption uses OpenRouter’s published token share.
How scores become comparable
Sources score on different scales (Elo, 0–100 indices, resolved-%), so each signal is standardized against the current board and mapped to 0–100: 50 is the board average and every 15 points is one standard deviation, clamped at the extremes. A model needs at least two independent evidence families to rank at all, and the board lists the top 25 by composite.
Per-use-case boards
The text, coding, vision, math, reasoning and agentic boards re-run the same blend over segment-appropriate signals — coding swaps in the coding benchmarks and SWE-bench Verified, vision uses the LMArena vision arena — and Blind Battles only ever pair models within the same segment.
Freshness and fair play
External feeds refresh daily and first-party telemetry is live; new models enter from external evidence within 48 hours of release. No self-reported vendor numbers are used anywhere, community buzz is shown as context but never feeds a rank, and every external source is attributed under the board.
Frequently asked questions
How is the AI model leaderboard ranked?
Each model gets one blended score from four evidence families — blind human preference (LMArena), independent benchmarks (Artificial Analysis, SWE-bench), OrcaRouter production reliability, and ecosystem adoption (OpenRouter) — combined with fixed public weights of 40 / 30 / 20 / 10.
Why does a brand-new model already have a rank?
New models enter the board from external evidence within 48 hours of release. The rank hardens as community battles and production traffic accumulate.
How often is the leaderboard updated?
External feeds refresh daily and OrcaRouter telemetry is live; the board shows its last-updated date in the header.
Why is a model missing from the board?
A model needs at least two independent evidence sources to rank, and the board lists the top 25 by composite score — models with thinner evidence stay unlisted until another source picks them up.
Are different use cases ranked separately?
Yes — the pills above the consolidated table switch between text, coding, vision, math, reasoning and agentic boards, and battles only pair like with like.
When does a new model’s rank harden?
Once a model carries three evidence families including OrcaRouter’s own battle and traffic data, its rank rests on the full blend rather than external evidence alone.
Where does the data come from?
LMArena (CC-BY-4.0), Artificial Analysis, SWE-bench and OpenRouter rankings, plus OrcaRouter production telemetry — all attributed under the board.
Can vendors game this ranking?
Hard to: the weights are public, no self-reported numbers are used, and blind community battles plus live traffic anchor every score.
Can I export or embed this data?
Yes — JSON and CSV export, an embeddable per-model rank badge, and an RSS feed of rank changes are free with attribution. Links are in the footer below.