
Reflection Beam, Explained: 501B Parameters, 23B Active, and Weights That Are Not Out Yet
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 145 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 128 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 933 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 183 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 217 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
On 5 October 2026 Reflection announced Reflection Beam, a sparse mixture-of-experts language model with 501 billion total parameters that activates 23 billion of them per token, and put a beta, OpenAI-compatible endpoint in front of it the same day. That is 4.6 per cent of the model awake at any moment — an aggressive ratio even by the standards of the current open-weight field — and it is the number the entire launch hangs on. Reflection says the full weights, a technical report and a model card are coming later this month under Apache 2.0. As of today they are not here, no price has been published anywhere on Reflection's own properties, and Artificial Analysis has no row for the model. So the honest summary of the launch is this: a very specific claim about efficiency, made by the lab that trained the model, with every supporting number supplied by that lab.
Reflection's own comparison tables place Reflection Beam against Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen3.8 Max and DeepSeek V4.1 Flash, which is a fair picture of the tier it is aiming at: strong-but-not-frontier open and semi-open models, several of them a fraction of Beam's size. Beam does not beat that field on raw capability, and we will show you exactly where it loses. What it claims to do is land near the top of it while spending roughly three to four times less inference compute than GLM 5.2 at comparable reasoning scores. That claim is the article.
What actually exists today
Everything in this section is from Reflection's own documentation, read on 6 October 2026. The API is live in beta behind a waitlist; the weights are not out.
• Model ID — Beam-501B-A23B, created 5 October 2026.
• Interface — an OpenAI-compatible surface at api.reflection.ai/openai/v1, exposing Chat Completions and the Models list, and nothing else. No Completions, no embeddings, no batches.
• Context — 256,000 tokens, with a footnote attached to the figure itself: "The context window may change during the beta." Maximum output is 128,000 tokens.
• Reasoning — a reasoning_effort parameter with five values: low, medium, high, xhigh and max. The default is medium, and the model always reasons regardless of the setting — there is no off switch.
• Modality — text in, text out. No image or audio input at launch, which is a real difference from Inkling, the multimodality-first model Mira Murati's Thinking Machines shipped in July.
• Knowledge cutoff — 30 June 2026.
• Price — not published. Reflection's blog post, its documentation and its model listing all omit it, and the platform console sits behind a signup wall.
That last line matters more than it looks. Every other model in Beam's comparison set has a public list price you can plug into a spreadsheet today. Beam does not, which means any cost argument about it right now is an argument about compute, not about dollars.

The 501-to-23 ratio is the argument
Sparsity is not new; the question is how far a lab is willing to push it. GLM 5.2 runs roughly 40 billion active parameters out of a total reported in the region of 744 billion — about 5.4 per cent. Reflection Beam sits at 4.6 per cent out of 501 billion. On paper that is a modest further step, and Reflection is careful about what it claims for it.
The efficiency figure in the launch post comes from a chart built on Artificial Analysis and DataCurve data, plotting reasoning score against estimated inference FLOPs, where FLOPs are approximated as twice the active parameter count times the mean number of generated tokens. That formula deliberately excludes prompt prefill, attention cost and serving overhead. It is a back-of-the-envelope model, not a measurement, and Reflection does not present it as one — but it is also the only quantitative support for the "3 to 4× cheaper at the same score" claim, and it was written by the vendor. Treat it as a hypothesis the weights, when they land, will let someone else test.
What the architecture buys in practice is a serving profile. Twenty-three billion active parameters is a model you can run on a comparatively small number of GPUs at interactive latency, and that is a different product from a 501-billion-parameter dense model even though the parameter count on the card is the same. It is the reason Reflection can talk about frontier-adjacent reasoning without talking about a data-centre-scale deployment.
Under four weeks to pre-train, and the disclosure is unusually specific
The training account is the part of the release that reads as genuinely informative, because Reflection published numbers most labs keep internal.
• Pre-training — 23.8 trillion tokens on 6,144 GB300 chips in NVL72 racks, finished in under four weeks at a reported 92.3 per cent goodput, with nine rewinds from checkpoints along the way.
• Reinforcement learning — 10,500 GB300 chips for four weeks, more than 100 million rollouts, maximum rollout context of 256,000 tokens, roughly 1.3 billion sandboxes and about one million distinct environments, with an average of 110,000 rollouts running concurrently.
• Midtraining — extends the model's effective context to 1 million tokens.
• Stability — asynchronous policy gradients that stay stable with day-scale staleness, plus a controllable length penalty so reasoning length can be tuned without retraining.
Read the pre-training and RL lines together and the shape of the claim appears. The efficiency story is not only about active parameters at inference; it is also about how fast the model was trained and how much of the compute went into environment-bound RL rather than into more tokens. A million environments and 1.3 billion sandboxes is a statement about tool-use and agentic training infrastructure, and it lines up with the benchmarks Reflection chose to lead with.
One number deserves a flag. The blog says midtraining extends the effective context to 1 million tokens; the API documentation lists a 256,000-token serving limit with a beta footnote attached. Those are a training-side capability and a serving-side limit, not a contradiction — but the gap is exactly the kind of thing that becomes a "1M context" headline if nobody reads the second document, and whoever writes about Beam next week should read the second document.

Benchmarks: where Reflection Beam wins, and where it does not
Every figure below is vendor-reported, from the tables in Reflection's launch post, and none of it has been independently reproduced. The comparison columns are the same numbers Reflection published for the other models; where a cell shows "NR" in the original, the model was not run. There is no Artificial Analysis row for Beam, so nothing here has been through a third-party harness.
Agentic coding
• Terminal-Bench v2.1 — Beam 80.1 vs GLM 5.2 81.0, GLM 5.3 88.2, Kimi K3 88.3, Qwen3.8 Max 86.6, DeepSeek V4.1 Flash 90.6, Inkling 63.8, Nemotron 3 Ultra 56.4.
• SWE-Bench Pro v1 — Beam 65.5 vs Qwen3.8 Max 67.7, GLM 5.2 62.1, Inkling 54.3, Nemotron 3 Ultra 46.4.
• SWE-Bench Pro v2-Hard — Beam 77.2 vs Kimi K3 88.2, GLM 5.3 84.3, Inkling 56.9.
• SWE-Bench Verified — Beam 80.9 vs Inkling 77.6, Nemotron 3 Ultra 70.7.
• SWE-Bench Multilingual — Beam 78.0 vs Nemotron 3 Ultra 67.7.
• DeepSWE v1.1 — Beam 44.4 vs DeepSeek V4.1 Flash 74.2, Kimi K3 68.0, GLM 5.3 61.0, Qwen3.8 Max 51.0, GLM 5.2 44.0.
• SWE Atlas Codebase QnA — Beam 34.6 vs Kimi K3 68.0, GLM 5.3 61.0.
That last row is the one to sit with. Long-horizon repository question answering is where a model has to hold a codebase in its head through a long interaction, and a 34.6 against 68.0 is not a rounding difference. DeepSWE tells a similar story: 44.4 puts Beam level with GLM 5.2 and far behind DeepSeek V4.1 Flash.
Reasoning
• AIME 2026 — Beam 97.8 vs GLM 5.2 99.2, Inkling 97.1.
• HLE, no tools — Beam 36.2 vs Kimi K3 46.9, Qwen3.8 Max 43.6, GLM 5.3 42.3, GLM 5.2 40.5, DeepSeek V4.1 Flash 39.1, Inkling 29.7, Nemotron 3 Ultra 26.7.
• GPQA Diamond — Beam 90.5 vs Kimi K3 93.5, Qwen3.8 Max 92.6, GLM 5.3 91.7, GLM 5.2 91.2, DeepSeek V4.1 Flash 90.9, Inkling 87.2, Nemotron 3 Ultra 87.
• SciCode — Beam 49.7 vs GLM 5.3 59.0, Kimi K3 58.7, Qwen3.8 Max 52.1, DeepSeek V4.1 Flash 52.0, Inkling 46.1, Nemotron 3 Ultra 44.6.
• CriPT AA — Beam 16.3 vs Kimi K3 23.4, GLM 5.2 20.9, Qwen3.8 Max 20.0, GLM 5.3 19.1, DeepSeek V4.1 Flash 14.3, Inkling 5.4, Nemotron 3 Ultra 3.1.
Beam's reasoning profile is mid-pack, and the pattern is consistent: comfortably ahead of the two models from the previous generation on Reflection's list, behind the current one. GPQA Diamond at 90.5 is a solid score that four other models beat.
Tool use and search
• MCP Atlas — Beam 78.7 vs Qwen3.8 Max 84.5, GLM 5.3 84.2, Kimi K3 82.3, GLM 5.2 77.8, Inkling 76.0, Nemotron 3 Ultra 63.1.
• AutomationBench, public split — Beam 37.0 vs DeepSeek V4.1 Flash 54.8, GLM 5.3 48.2, Kimi K3 46.7, Qwen3.8 Max 39.8, GLM 5.2 26.2.
• tau3 banking — Beam 38.0 vs Qwen3.8 Max 55.2, GLM 5.2 37.1, Kimi K3 37.1, Inkling 25.0, Nemotron 3 Ultra 22.6.
• BrowseComp with context management — Beam 77.4 vs Kimi K3 91.2, Inkling 77.1, Nemotron 3 Ultra 44.4.
• DeepSearchQA with context management — Beam 80.1 vs Kimi K3 95.0.
Reflection also showed two capability demos in the post: a 95.5 per cent solve rate on the "Land or Water" puzzle, a 180-by-90 grid with 16,200 points that requires holding a large board state — a figure that sits between the 92.5 and 97.8 per cent Reflection attributes to Opus 5 and Fable 5 on the same task — and a Text2SQL fine-tune of the smallest Gemma-4 that raised held-out accuracy to 66.5 per cent. Demos are curated; they show the model can do a thing, not how often.
What you can do with it today, and what you should not plan around
Today, exactly one thing is generally possible: request beta access and call Beam-501B-A23B through Reflection's own endpoint with an OpenAI-shaped client. There is no published price, so there is no way to model the cost of a workload on it. There are no weights, so there is no self-hosting, no quantisation, no fine-tune and no third-party reproducibility. There is no independent benchmark row, so the entire performance picture above rests on one lab's tables.
Reflection Beam is not one of our routes. We are not going to imply otherwise, and we are not going to point you at a reseller who might have it. What we can tell you is what the comparison set costs, because we route four of those models and the figures below are read live off our own catalogue this morning: GLM 5.2 at $1.40 in and $4.40 out per million tokens, GLM 5.3 at $1.26 and $3.96, Qwen3.8 Max at $2.00 and $6.00, and DeepSeek V4.1 Flash at $0.15 and $0.60. That is a real spread — DeepSeek V4.1 Flash is nearly ten times cheaper on input than GLM 5.2 — and it is the reason a beta model with no price is hard to slot into a decision. OrcaRouter runs one OpenAI-compatible key across those and 200-plus other models with the provider's list price passed through and nothing added, which means a vendor price change is live on our side the same day rather than whenever a reseller refreshes. And because automatic failover is part of the routing rather than something you build, a new and unproven model is the kind of thing you can put on a share of traffic instead of on your critical path — which is the only responsible way to use something whose benchmarks nobody has reproduced yet.

What to watch before you take the efficiency claim seriously
Three things will settle the question, and two of them are promised within the month.
The first is the weights. Apache 2.0 and a technical report would let anyone with a couple of nodes measure the thing that actually matters — tokens per second per GPU at a fixed quality bar, and whether 23 billion active parameters really does deliver the score-per-FLOP the chart implies. Until then the efficiency argument is arithmetic on a napkin.
The second is a price. A model with no list price has no place in a cost comparison, and if Beam lands above the GLM and DeepSeek tier, the compute story stops mattering for anyone with a budget. If it lands below, the picture changes quickly.
The third is a third party running the suite. Every number above comes from the same document, and a vendor-reported table that puts its own model behind on seven of sixteen rows is a good sign about the lab's honesty — but it is still one lab, one harness, one set of prompts.
If you need a capable agentic coder today and you need to know what it costs, the answer is not Reflection Beam yet. It might be in three weeks. The model an efficiency claim is worth betting on is the one whose weights you can download and whose FLOPs you can count yourself — and on that test, Reflection has given us a date and not a model.
Compared in this article4
Detected from this article · Benchmarks: Artificial Analysis · updated daily
