
Reflection Beam vs Kimi K3: The Same Benchmark Name, Two Different Harnesses
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 151 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 104 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 248 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 232 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Reflection Beam scores 80.1 on Terminal-Bench 2.1. Kimi K3 scores 88.3. Put those two figures side by side, as most of this week's coverage has, and you conclude Beam is roughly eight points off the best open model in the field. But those two numbers do not come from the same room. Beam's 80.1 is vendor-reported, taken from the table in Reflection's own launch post. Kimi K3's 88.3 is also vendor-reported, this time from Moonshot's own table — and a vendor's table prints its competitor's score only when it has been run on the vendor's own harness, under the vendor's own prompt set, at the vendor's own effort setting. A third party, Artificial Analysis, has since run Kimi K3 on a benchmark with the same name and published 85.0. That is a 4.9-point gap, not 8.2 — and the direction of the correction is entirely predictable, because the number being eroded is always the one that came from the lab with something to prove.
That is the whole problem with the comparison this article's title promises, and it is the reason this piece spends its first half on provenance rather than on scores. Reflection Beam is a 501-billion-parameter sparse mixture-of-experts model with 23 billion parameters active per token, announced 5 October 2026 by Reflection AI with a beta OpenAI-compatible endpoint live the same day. Kimi K3 is Moonshot AI's flagship open model, listed in our own catalogue with a release date of 15 July 2026 and independently scored by Artificial Analysis. One of them is a shipping product with a rate card and a third-party harness run; the other is a waitlisted beta whose weights, technical report and price do not exist yet. Comparing them is not a symmetric exercise, and pretending otherwise would produce a page that looks precise and is wrong.
The 30-second version
• Beam is faster per FLOP on paper, unpriced in practice, and unreproduced. Kimi K3 is scored by a third party, priced at $3.00 input and $15.00 output per million tokens, and generally available.
• On the one benchmark both have been run on — Terminal-Bench 2.1 — the honest gap is about five points, not eight, once each side's numbers come from a neutral harness.
• Beam's real advantages are structural, not numeric: a 23B active-parameter serving profile and a promised Apache 2.0 licence. Neither is usable today.
• If you need a model this week, this is not a coin flip. It is one shipping model and one waitlist.
What Reflection actually published, and against whom
Reflection's launch post runs one scorecard against seven other models: Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8 Max and DeepSeek V4.1 Flash. Every figure in the table is vendor-reported, none of it has been independently reproduced, and there is no Artificial Analysis page for Reflection Beam — the model page 404s on their site as of 6 October 2026. Several cells in the original are marked NR because the model was not run.
Here is the entire Beam-versus-K3 slice of that table, one line per dimension:
• Terminal-Bench 2.1 — Beam 80.1 vs Kimi K3 88.3
• SWE-Bench Pro v2-Hard — Beam 77.2 vs Kimi K3 88.2
• DeepSWE v1.1 — Beam 44.4 vs Kimi K3 68.0
• SWE Atlas Codebase QnA — Beam 34.6 vs Kimi K3 68.0
• HLE, no tools — Beam 36.2 vs Kimi K3 46.9
• GPQA Diamond — Beam 90.5 vs Kimi K3 93.5
• SciCode — Beam 49.7 vs Kimi K3 58.7
• MCP Atlas — Beam 78.7 vs Kimi K3 82.3
• BrowseComp with context management — Beam 77.4 vs Kimi K3 91.2
• DeepSearchQA with context management — Beam 80.1 vs Kimi K3 95.0
Read that list honestly and the shape of the launch appears. Reflection is not claiming a win over K3 anywhere. It is claiming to land near the top of a tier while spending three to four times less inference compute than GLM 5.2 at comparable reasoning scores — and the one row where the two models are close, Terminal-Bench 2.1, is the row where the comparison is least trustworthy.
Why the harness matters more than the score
A benchmark name is a label, not a specification. Terminal-Bench 2.1 means a particular set of tasks run under a particular agent scaffold, with a particular retry policy, a particular token budget and a particular reasoning-effort setting. Two labs can both honestly report a "Terminal-Bench 2.1" figure and produce numbers that are not comparable, because the scaffold and the effort setting are where most of the variance lives. Artificial Analysis exists as an organisation precisely because of this: it runs one harness, in one configuration, across many models, so that its numbers can be subtracted from each other.
So the 80.1 that Reflection prints for Beam is a vendor number on a vendor harness, and the 88.3 it prints for K3 is also a vendor number — the vendor being Reflection, measuring its own competitor. That second figure is the least reliable number in the row: a lab running a rival's weights on its own scaffold has no incentive to run them at the rival's best configuration, and no obligation to disclose the one it chose. There is nothing dishonest in it; it is simply a number whose provenance does not support the weight the coverage has put on it.
Artificial Analysis's own run has Kimi K3 at 85.02 on Terminal-Bench 2.1, alongside a 12.6 on Terminal-Bench v4.0 — a different and harder test — an AA Coding Index of 76.2 (rank 9 of the models on the board), an Intelligence Index of 43.6, GPQA Diamond 93.5, HLE 46.9, SciCode 59.5, tau-banking 45.98 and a Long-Context Recall of 88.7. Those are read off K3's catalogue card in our own system, sourced to artificialanalysis.ai, with the evaluation snapshot dated 15 July 2026. Beam has one number in that list's universe and it comes from Reflection. That absence should be temporary rather than permanent: Artificial Analysis has said Reflection gave it access and that it is benchmarking the model, and its own early wording — that Beam "will be one of the most token-efficient open models we've seen for its level of intelligence" — is a benchmark house signalling an expectation, not reporting a result. Until that score prints, this article's numbers are not subtractable, and we are not going to pretend otherwise.

What each one costs, and what each one is
The cost comparison is not close, and it is not really a comparison at all, because one side of it is blank.
• Price — Kimi K3 $3.00 input / $15.00 output per million tokens, with cache reads at $0.30, vs Reflection Beam: no published price anywhere on Reflection's blog, documentation or model listing, with the platform console behind a signup wall.
• Context — 1,048,576 tokens on Kimi K3 vs 256,000 on the Beam beta, with a footnote on Beam's own documentation reading "the context window may change during the beta."
• Modality — Kimi K3 accepts text and images; Reflection Beam is text-in, text-out with no image or audio path at launch.
• Availability — Kimi K3 is generally available today; Reflection Beam is a waitlisted beta with weights promised for later in October under Apache 2.0.
• Independent scores — K3 has a full Artificial Analysis row; Beam has none.
• Serving profile — Kimi K3 is a large dense-ish frontier model at 46.8 output tokens per second on our own playground measurement over the past week; Beam's 23B active parameters are a genuine architectural argument for lower latency and lower cost per token, but there is no endpoint open to the public to measure it on.
The efficiency claim deserves one more sentence of care, because it is the launch's headline and it is doing more work than it can bear. Reflection's "3 to 4× less inference compute" comes from a chart that approximates FLOPs as twice the active parameter count times the mean number of generated tokens. That formula deliberately excludes prompt prefill, attention cost and serving overhead — the parts of the bill that dominate long-context agentic work. It is a back-of-the-envelope estimate of a compute ratio, not a measurement, not a dollar figure, and it was built by the vendor. Treat it as a hypothesis the weights will let someone else test.
Where OrcaRouter sits in this
Kimi K3 is one of our routes. It is listed at $3.00 input and $15.00 output per million tokens with cache reads at $0.30, which is Moonshot's provider rate passed through with nothing added — so if Moonshot moves its rate card, the figure on our page moves the same day rather than whenever a reseller refreshes. It is callable from one OpenAI-compatible key alongside 200-plus other models, which matters here for a specific reason: the honest answer to "Beam or K3?" is that a team hedging should not choose. Because automatic failover is part of the routing rather than something you build, a model like Beam — beta, unpriced, unscored by any third party — is the exact thing you put on a share of traffic rather than on your critical path. You route the work to K3 by default, send Beam a slice when your waitlist invite lands, and let the routing fall back to K3 on error. That is a decision you can make about an unproven model. Betting a production path on one is not.

What would change this verdict
Three things, in descending order of how much they matter.
A price. Beam landing below the K3 tier would make the efficiency argument concrete instead of theoretical. Second, the weights. Apache 2.0 plus a technical report would let anyone with a couple of nodes measure tokens per second per GPU at a fixed quality bar — the test that actually settles a compute claim. Third, a third-party harness run. Reflection has given Artificial Analysis access and a score is expected out of it; until that number prints, every Beam-versus-K3 figure in circulation traces back to a document written by one of the two vendors.
Two questions worth answering directly
Is Reflection Beam better than Kimi K3?
On the evidence available today, no — and nobody has published evidence that it is. Reflection's own table has K3 ahead on all ten shared rows above, several of them by margins (SWE Atlas 34.6 vs 68.0, DeepSearchQA 80.1 vs 95.0) that are not close. Beam's case is cost per unit of capability, and that case is a vendor-drawn chart until the weights and a price exist.
Should I wait for Beam's weights before building on K3?
Only if self-hosting is a requirement rather than a preference, because that is the one axis where the two are not substitutes — K3 is API-only while Beam promises Apache 2.0 weights. If you are calling an API, K3 is available, scored and priced, and Beam is a waitlist. There is no version of that decision worth delaying a project for.

Reflection has given us a date and a chart, and a date is not a model. The one thing the launch genuinely establishes is that a 501B model activating 23B parameters can be trained to sit within a few points of the open frontier — which, if the weights arrive and the numbers hold, is a meaningful result about efficiency. It is not yet a reason to move work off a model that a third party has actually measured.
