
Reflection Beam Ships an API in Beta, and the Weights Are Still Two Weeks Out
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 150 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 127 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 53 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 204 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 229 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Reflection's first model is called Reflection Beam, and the interesting thing about it this morning is not that it exists — it is that only half of it does. The company announced Beam on 5 October 2026 as a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, and put a beta, OpenAI-compatible endpoint in front of it the same day. The weights are promised for later this month under Apache 2.0, alongside a technical report, a model card and the serving stack. Until they land, the only people who can run Beam-501B-A23B are the ones Reflection hands a waitlist key to, and every performance figure in circulation comes from one lab's own tables.
That is a strange place for a launch to stop, and it is worth being precise about which half we have. Reflection has shipped a preview you can sign up for and a set of benchmark tables nobody has reproduced. It has not shipped the thing "open-weight" in the headline refers to.
What is actually live this morning
Everything in this section is from Reflection's own blog post and documentation, read on 6 October 2026.
• Model ID — Beam-501B-A23B, created 5 October 2026, served at api.reflection.ai/openai/v1 with Chat Completions and a Models list and nothing else.
• Shape — 501 billion total parameters, 23 billion active per token, giving an activation ratio of about 4.6 per cent. For scale, GLM 5.2 sits near 5.4 per cent.
• Context — 256,000 tokens on the API, with a footnote attached to the figure itself warning that the window may change during the beta. Maximum output is 128,000 tokens.
• Reasoning — a reasoning_effort parameter with five settings: low, medium, high, xhigh and max. Medium is the default, and the model always reasons. There is no off switch and no non-reasoning mode.
• Modality — text in, text out. No image, audio or video input at launch.
• Price — not published. The blog post does not carry one, the documentation does not carry one, and the platform console sits behind a signup wall.
• Access — a waitlist. There is no self-serve path and there are no weights, so there is no download, no quantisation and no fine-tune.
The price line is the one that changes how you read everything else. Four of the seven models Beam is compared against in Reflection's own tables have public list prices you can put in a spreadsheet today. Beam does not, so any cost claim about it right now is a claim about compute, not about dollars.

The number the launch hangs on
Reflection's pitch is inference efficiency, and it is unusually specific about what that means: on advanced reasoning benchmarks, Beam scores comparably to GLM 5.2 while using 3 to 4× less inference compute. The gains are described as larger still against models in the 2-trillion-parameter family, where Qwen 3.8-Max sits.
The chart behind that claim is worth reading carefully, because it is the load-bearing evidence in the whole post. It plots reasoning score against estimated inference FLOPs across DeepSWE, Humanity's Last Exam and Terminal-Bench 2.1, and it estimates FLOPs as roughly twice the active parameter count times the mean number of generated tokens per attempt. That formula deliberately excludes prompt prefill, context-dependent attention operations and serving overhead — Reflection says so in the caption. It is a consistent comparison between models rather than a measurement of anybody's bill, and it is also the only quantitative support for the efficiency claim, written by the lab that built the model. Treat it as a hypothesis the weights will let somebody else test.
What the architecture does buy in practice is a serving profile. Twenty-three billion active parameters is a model you can plausibly run on a comparatively small number of GPUs at interactive latency, which is a different product from a 501-billion-parameter dense model even though the number on the card is the same. That is why Reflection can talk about frontier-adjacent reasoning without talking about a data-centre-scale deployment — and why the eventual weight release is the event that matters more than the beta key.
Where Beam lands on Reflection's own tables
Every figure below is vendor-reported, from the tables in Reflection's launch post, and none of it has been independently reproduced. Where Reflection published no number for a model, the cell reads NR in the original. The comparison columns are Reflection's, not ours and not a third party's.
Coding and terminal work
• Terminal-Bench v2.1 — Reflection Beam 80.1 against GLM 5.2 81.0, GLM 5.3 88.2, Kimi K3 88.3, Qwen 3.8-Max 86.6 and DeepSeek V4.1 Flash 90.6. Beam sits below every one of them.
• SWE-Bench Verified — Reflection Beam 80.9 against Inkling 77.6 and Nemotron 3 Ultra 70.7. This is where Beam looks strongest, but note that only the two weakest models on the list were run on it.
• SWE-Bench Pro v1 — Reflection Beam 65.5 against Qwen 3.8-Max 67.7, GLM 5.2 62.1, Inkling 54.3, Nemotron 3 Ultra 46.4.
• SWE-Bench Pro v2-Hard — Reflection Beam 77.2 against Kimi K3 88.2, GLM 5.3 84.3, Inkling 56.9. A ten-point gap to the top of the table.
• DeepSWE v1.1 — Reflection Beam 44.4 against DeepSeek V4.1 Flash 74.2, Kimi K3 68.0, GLM 5.3 61.0, Qwen 3.8-Max 51.0, GLM 5.2 44.0. Beam lands level with the previous generation and thirty points behind the leader.
• SWE Atlas Codebase QnA — Reflection Beam 34.6 against Kimi K3 68.0 and GLM 5.3 61.0. Holding a large repository in mind across a long interaction is the hardest thing on this list, and Beam's 34.6 is not a rounding difference.
Reasoning and science
• AIME 2026 — Reflection Beam 97.8 against GLM 5.2 99.2 and Inkling 97.1.
• HLE without tools — Reflection Beam 36.2 against Kimi K3 46.9, Qwen 3.8-Max 43.6, GLM 5.3 42.3, GLM 5.2 40.5, DeepSeek V4.1 Flash 39.1, Inkling 29.7, Nemotron 3 Ultra 26.7. Mid-pack, and ahead of only the two previous-generation models.
• GPQA Diamond — Reflection Beam 90.5 against Kimi K3 93.5, Qwen 3.8-Max 92.6, GLM 5.3 91.7, GLM 5.2 91.2, DeepSeek V4.1 Flash 90.9.
• SciCode — Reflection Beam 49.7 against GLM 5.3 59.0, Kimi K3 58.7, Qwen 3.8-Max 52.1, DeepSeek V4.1 Flash 52.0.
• CriPT AA — Reflection Beam 16.3 against Kimi K3 23.4, GLM 5.2 20.9, Qwen 3.8-Max 20.0, GLM 5.3 19.1, DeepSeek V4.1 Flash 14.3, Inkling 5.4, Nemotron 3 Ultra 3.1.
Tools, search and long context
• MCP Atlas — Reflection Beam 78.7 against Qwen 3.8-Max 84.5, GLM 5.3 84.2, Kimi K3 82.3, GLM 5.2 77.8.
• AutomationBench, public split — Reflection Beam 37.0 against DeepSeek V4.1 Flash 54.8, GLM 5.3 48.2, Kimi K3 46.7, Qwen 3.8-Max 39.8, GLM 5.2 26.2.
• tau3 banking — Reflection Beam 38.0 against Qwen 3.8-Max 55.2, GLM 5.2 37.1, Kimi K3 37.1.
• BrowseComp with context management — Reflection Beam 77.4 against Kimi K3 91.2, Inkling 77.1, Nemotron 3 Ultra 44.4.
• DeepSearchQA with context management — Reflection Beam 80.1 against Kimi K3 95.0.
• AA-LCR — Reflection Beam 79.3 against Kimi K3 88.7, DeepSeek V4.1 Flash 84.0, Qwen 3.8-Max 80.3, GLM 5.3 79.7, Nemotron 3 Ultra 79.3, GLM 5.2 78.3, Inkling 77.3. Long-context recall is the one general-capability row where Beam is genuinely competitive rather than mid-pack.
Read the four tables together and a consistent shape appears: Beam beats the two previous-generation models on Reflection's list and loses to the current one, on almost every row, by margins that range from small to disqualifying. A vendor publishing tables that put its own model behind on that many rows is a good sign about the lab's honesty. It is not a sign that the model is at the frontier.

The one independent voice so far, and what it does not say
The only third-party assessment on the record is Artificial Analysis, which said on 5 October that Reflection had given it access and that it was independently benchmarking Beam, adding that early indicators suggest Beam will be one of the most token-efficient open models it has seen for its level of intelligence.
That is a meaningful statement from the organisation whose index is the one most buyers actually read, and it is also a statement without a number in it. As of this morning there is no Beam row on Artificial Analysis's leaderboard — the model pages return 404 and the leaderboard payload carries no Beam entry — so the token-efficiency claim is, for now, a promise from a third party that it will publish something. Nothing in the index has been recomputed on Beam yet.
The training disclosure, which is the strongest part of the release
The engineering section of Reflection's post contains numbers most labs keep internal, and they are worth recording because they are the part of the release that will still be interesting in a month.
• Pre-training — 23.8 trillion curated tokens on 6,144 NVIDIA GB300 chips in NVL72 racks, finished end to end in under four weeks, at a reported 92.3 per cent goodput, with nine semi-automatic rewinds along the way attributed to gradient-norm spikes or suspected silent data corruption.
• Reinforcement learning — 10,500 GB300 chips for four weeks, more than 100 million rollouts, a maximum rollout context of 256,000 tokens, roughly 1.3 billion sandboxes and about one million distinct environments sourced for coding, agentic and STEM tasks.
• Serving — new weights reached the inference fleet in a median of about 12 seconds, with hierarchical distribution cutting cross-rack traffic by 75 per cent and making fleet-wide adoption 2.2× faster than every replica pulling weights itself.
• Stability — fully asynchronous policy gradients that stay numerically stable when learning from interactions a day stale, plus a controllable length penalty for trading reasoning length against score without retraining.
• Data — about 95 per cent of raw internet tokens eliminated through parsing, deduplication and curation, with the claim that conventional filters would have missed roughly 1.8 trillion of the tokens Reflection retained, including 87 per cent of its curated web-code tokens.
Two lines deserve a flag. The post says midtraining extends Beam's effective context to 1 million tokens, while the API documentation lists a 256,000-token serving limit with a beta footnote. Those are a training-side capability and a serving-side limit rather than a contradiction, but the gap is exactly what becomes a "1M context" headline if nobody reads the second document. And the RL figure of more than 100 million rollouts is presented as one of the largest such runs by any open lab, which is a claim about scale, not about what the scale bought — the score-versus-rollouts curve is the vendor's own and stops where the vendor stopped plotting it.

What you can do with Reflection Beam today
One thing: join the waitlist and, if you get a key, call Beam-501B-A23B through Reflection's own endpoint with an OpenAI-shaped client. There is no published price, so there is no way to model a workload's cost on it. There are no weights, so there is no self-hosting, no quantisation and no third-party reproduction. There is no independent benchmark row, so the entire performance picture above rests on a single lab's tables.
Reflection Beam is not one of our routes, and we are not going to imply that it is. What we can give you is the price of the field it is trying to enter, read live off our own catalogue this morning: GLM 5.2 at $1.40 in and $4.40 out per million tokens, GLM 5.3 at $1.26 and $3.96, Qwen 3.8-Max at $2.00 and $6.00, and DeepSeek V4.1 Flash at $0.15 and $0.60. That is a spread of more than nine times between the cheapest and the dearest on input alone — and it is why a beta model with no list price is genuinely hard to slot into a decision, no matter how good its efficiency chart looks.
OrcaRouter runs one OpenAI-compatible key across those four and 200-plus other models, passing the provider's list price through with nothing added, so a vendor price change is live on our side the same day rather than whenever a reseller refreshes a table. Automatic failover is part of the routing rather than something you build around each provider, which means an unproven model — with no reproduction, no independent score and no price — is exactly the kind of thing you put on a share of traffic instead of on your critical path.
When the claim becomes testable
Three things settle this, and Reflection has put two of them inside the month.
The first is the weights. Apache 2.0 plus a technical report lets anyone with a couple of nodes measure the thing that matters — tokens per second per GPU at a fixed quality bar, and whether 23 billion active parameters really delivers the score-per-FLOP the chart implies. Until then the efficiency argument is arithmetic on a napkin, and everyone repeating it is repeating Reflection's caption.
The second is a price. A model with no list price has no place in a cost comparison. If Beam lands above the GLM and DeepSeek tier, the compute story stops mattering for anyone with a budget; if it lands below, the picture changes quickly.
The third is the index. Artificial Analysis has said it is benchmarking Beam and has published no number. When that row appears, it will be the first figure in this story that Reflection did not compute.
If you need a capable agentic coder today and you need to know what it costs, the answer is not Reflection Beam yet. It might be in three weeks. The model an efficiency claim is worth betting on is the one whose weights you can download and whose FLOPs you can count yourself — and on that test, Reflection has given us a date and a waitlist, not a model.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
