A generated hero card for the article 'Reflection Beam vs Gemini 3.1 Pro Preview: One Reads Video, One Reads Text', titled 'Reflection Beam vs Gemini 3.1 Pro Preview' with the subtitle 'One reads video, one reads text' and three chips reading 'Text only vs five modalities', '256K vs 1M context' and 'No price vs a tiered rate'.
Guides & Insights

Reflection Beam vs Gemini 3.1 Pro Preview: One Reads Video, One Reads Text

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Start with the asymmetry that no benchmark table in this matchup mentions. Reflection Beam, announced on 5 October 2026, is a 501-billion-parameter sparse mixture-of-experts model with 23 billion parameters active per token, and it takes text in and returns text out. Nothing else. Gemini 3.1 Pro Preview, the vendor's frontier multimodal model since 19 February 2026, accepts text, images, video, audio and files, and it is the only model in this comparison where the question "can it see the thing I want to ask about" has a different answer for each side. Everything else — the sparsity ratios, the context windows, the price tiers — is an argument. That one is a specification.

It also means this is the least symmetric pair in the set, and the most useful, because it exposes what a model comparison is actually for. If your workload is text, Beam and Gemini 3.1 Pro are two options on a spectrum that runs from cheap-and-unmeasured to expensive-and-thoroughly-scored. If your workload has a frame of video in it, there is no comparison to make.

What each model is

Gemini 3.1 Pro Preview is Google's preview-tier flagship: 1,048,576 tokens of context, a maximum output of 65,536 tokens, five input modalities, and a 1-million-token window reached through a price ladder rather than a flat rate. Google positions it for enhanced software engineering, improved agentic reliability and more efficient token usage on complex workflows. It is not open weights. Its name still carries "Preview", which is a status Google has held for this model since February.

Reflection Beam is Reflection's first model and the beginning of an open-weight series. Five hundred and one billion total parameters, 23 billion active — about 4.6 per cent of the model awake per token. Twenty-three point eight trillion pre-training tokens on 6,144 GB300 chips in under four weeks, then a reinforcement-learning campaign on 10,500 GB300 chips for four weeks with more than 100 million rollouts across roughly one million environments. Its API is OpenAI-compatible at api.reflection.ai/openai/v1 with Chat Completions and a Models list, its context is 256,000 tokens with a beta footnote, its reasoning_effort parameter has five levels and no off switch, and its weights are promised for later this month under Apache 2.0.

• Modality — Gemini 3.1 Pro Preview takes text, image, video, audio and file; Reflection Beam takes text only.

• Context and output — 1,048,576 tokens in and 65,536 out for Gemini, against 256,000 in and 128,000 out for Beam.

• Price — Gemini 3.1 Pro Preview at $2.00 in and $12.00 out per million tokens up to 200,000 tokens, rising to $4.00 and $18.00 above that, with cache reads at $0.20; Reflection Beam has no published price.

• Tooling surface — native tool calling, structured outputs and search grounding on the Google side; tool calling and structured outputs on the Beam side, with an endpoint limited to Chat Completions.

• Weights — proprietary for Gemini; Apache 2.0 promised for Beam, not yet released.

• Independent score — Gemini 3.1 Pro Preview carries an Artificial Analysis index; Beam has no row on the board.

The modality gap is the whole argument

It is worth being concrete rather than gesturing at "multimodal," because the practical consequences are specific.

A workload that ingests a screen recording, a product photo, a scanned contract or a call centre audio file cannot be served by Beam at any price, with any prompt, on any plan. There is no workaround at the API layer: Beam's documentation describes one input modality and one output modality, and no image or audio pathway is listed. Gemini 3.1 Pro Preview handles all five, which means it is the only model here that can be dropped into a workflow where the input is not already text.

The reverse case is worth stating too, because it is the one people get wrong. If your input is text and will stay text, Gemini's five modalities buy you nothing, and you are paying a multimodal model's price to run a text workload. That is not an argument for Beam — it is an argument that the modality question should be answered before the benchmark question, because it eliminates options cheaper and more reliably than any score comparison does.

Two previews, two different meanings

Both model names carry a caveat and the caveats are not alike, which is the most interesting thing about this pairing.

Gemini 3.1 Pro's "Preview" is a naming convention on a model that has been generally callable since February, with an independent score, a published rate card including a documented long-context tier, and a full tool surface. In practice "preview" here means Google reserves the right to supersede it, not that it is hard to get or unscored.

Beam's beta is a genuine gate. Access is a waitlist, the model ID was created on 5 October 2026, there is no rate card, no third-party score, and the artefacts that would let anyone verify the central claim — weights, technical report, model card — are promised rather than shipped. The context figure carries a footnote warning it may change during the beta, which is a hedge no generally available model needs.

The consequence is asymmetric in a way that matters for procurement. A team adopting Gemini 3.1 Pro Preview is accepting model churn risk. A team adopting Beam today is accepting an unknown price, an unknown independent score, and no ability to run the model themselves. Those are different categories of risk, and price is the one that most often decides.

A generated two-column scoreboard titled 'Reflection Beam vs Gemini 3.1 Pro Preview — the scoreboard'. Left column 'Reflection Beam': Input modalities text only; Context 256,000 tokens (beta); Max output 128,000 tokens; Price not published; Availability waitlisted beta; Independent score none yet. Right column 'Gemini 3.1 Pro Preview': Input modalities text, image, video, audio, file; Context 1,048,576 tokens; Max output 65,536 tokens; Price $2.00 / $12.00 to 200K, then $4.00 / $18.00; Availability preview, callable since February 2026; Independent score AA index 29.7. Footer reads 'Beam figures vendor-reported and unaudited. Gemini price is the provider list rate; its index is Artificial Analysis, read 6 October 2026.'

Benchmarks, and the citing problem

Reflection's launch tables do not include Gemini 3.1 Pro Preview or any Google model. Its comparison set is open and semi-open only: Inkling, Nemotron 3 Ultra, GLM 5.2, GLM 5.3, Kimi K3, Qwen 3.8-Max and DeepSeek V4.1 Flash. So the two models have never been run against each other in a published table, and the figures below come from different harnesses on both sides.

• Artificial Analysis Intelligence Index — Gemini 3.1 Pro Preview records 29.7, ranked 58th of 147 in its run. Beam has no index row at all.

• GPQA Diamond — Gemini 3.1 Pro Preview at 94.1 per Artificial Analysis against Beam's vendor-reported 90.5.

• SciCode — 58.7 for Gemini against Beam's vendor-reported 49.7.

• Humanity's Last Exam — 47.0 for Gemini against Beam's vendor-reported 36.2, the latter explicitly without tools.

• Terminal-Bench v2.1 — 73.8 for Gemini per Artificial Analysis against Beam's vendor-reported 80.1. This is Beam's strongest row in the matchup, and it is a harvest against a model that has been on the market since February.

• Long-context recall — 82.0 for Gemini against Beam's vendor-reported 79.3 on AA-LCR.

• tau-squared Bench — 95.6 for Gemini against Beam's 78.7 on MCP Atlas and 38.0 on tau3 banking. Related evaluations of agentic tool use, not the same one, and the naming difference matters.

There is a separate honesty problem with the Gemini side of this table that deserves its own sentence. Independent index scores for Gemini 3.1 Pro Preview have been cited at 30, 46, 47.7 and 57 across different sources, and Artificial Analysis serves different index revisions on different model pages at the same moment. The 29.7 figure above is the one on the page we read on 6 October 2026, and it is the only one we will quote — but any article that puts a single Gemini index number next to a competitor without saying which snapshot it came from is presenting a floating number as a fixed one. The same caution applies in reverse to Beam, where there is no snapshot at all.

Price, and the tier nobody plans for

Gemini 3.1 Pro Preview is $2.00 in and $12.00 out per million tokens up to 200,000 tokens, and $4.00 and $18.00 above it. Cache reads are $0.20 per million and cache writes $0.375. The tier boundary is the part worth planning around: the model's headline selling point is a 1,048,576-token window, and the moment a request crosses 200,000 tokens the input rate doubles and the output rate rises by half. A workload designed around the million-token window is therefore a workload priced at the second tier for most of its traffic, not the first.

Beam has no rate card, so there is no tier to model and nothing to compare. That absence is not neutral: it means the cheapest defensible statement about Beam's cost is that it is unknown, and the only quantitative support for its efficiency positioning is Reflection's own chart, which estimates generated FLOPs as twice the active parameter count times mean generated tokens per attempt across DeepSWE, HLE and Terminal-Bench 2.1 — with a caption conceding that prompt prefill, attention and serving overhead are excluded. On the workloads where a million-token context matters, prefill is not a rounding error, and it is precisely the term the formula leaves out.

A screenshot of the OrcaRouter model page for Gemini 3.1 Pro Preview (google/gemini-3.1-pro-preview), captured 6 October 2026, showing the model name and id, the INPUT $2.00 and OUTPUT $12.00 pricing tiles, p50 TTFT 10.00 s, p95 TTFT 10.00 s, 13.9M tokens of 7-day traffic, the provenance lines 'Public benchmarks by Google - 2026-02-19' and 'AI Analysis', the summary describing a frontier reasoning model with improved agentic reliability, and a code sample using base_url https://api.orcarouter.ai/v1.

Tool use, search and where the two designs differ

Gemini 3.1 Pro Preview's advertised strengths are software engineering, agentic reliability and token efficiency, and its published tool surface includes function calling, structured outputs and search grounding. That last item is the one to note for anyone comparing agentic stacks: grounded search is a first-party capability on Google's side, and it changes what a tool-using agent can do without an external retrieval service wired in.

Beam's tool story is more interesting than a spec line suggests and harder to evaluate. Reflection reports MCP Atlas at 78.7, AutomationBench public at 37.0, tau3 banking at 38.0, BrowseComp with context management at 77.4 and DeepSearchQA with context management at 80.1 — and notes that during reinforcement learning the model organically learned to query other language models and to call OCR APIs when given web access, despite browsing tasks being absent from the training mixture. If that generalisation holds, it is a real signal about agentic transfer. It is also, at this point, a vendor observation from a training run, with no independent check.

On the tool-use rows where both sides have numbers, the picture is mixed rather than one-sided. Beam's vendor table puts it ahead of GLM 5.2 on MCP Atlas and level with GLM 5.2 and Kimi K3 on tau3 banking, while Gemini's tau-squared Bench figure of 95.6 is the highest agentic number either side has published. Different suites, different harnesses, and no shared evaluation — which is the recurring problem in this comparison.

Choosing, and how to hedge

The decision rule that falls out of the above is simple enough to state in one line: if any input is not text, Beam is not an option and the comparison is over. If every input is text, the question becomes whether Beam's unattributed efficiency claim is worth an unknown price and no independent score, set against a model that costs $2.00 and $12.00 with a documented ladder and a full tool surface.

Gemini 3.1 Pro Preview is on our catalogue, with the provider's list rate passed through and nothing added, behind one OpenAI-compatible key at api.orcarouter.ai/v1 alongside 200-plus other models — and the same key serves the models Beam competes with on price, from GLM 5.3 at $1.26 and $3.96 to DeepSeek V4.1 Flash at $0.15 and $0.60, read live this morning. Because the tier boundary at 200,000 tokens is a routing problem rather than a model problem, the routing DSL lets you express it directly: keep short requests on the cheap tier and pin the genuinely long ones where the window is the reason you chose the model, in one call rather than in application code.

Reflection Beam is not one of our routes, and we will not point you at anybody who claims to have it. If the Apache 2.0 weights arrive this month as promised, self-hosting becomes a third option for text-only work and the efficiency claim becomes testable on your own hardware.

A generated single-column card titled 'Reflection Beam — the state of play, 6 October 2026', used as the article's closing fact sheet. Rows read 'Total / active parameters: 501B / 23B', 'Pre-training: 23.8T tokens on 6,144 GB300, under four weeks', 'RL campaign: 10.5K GB300, 4 weeks, 100M+ rollouts', 'API context: 256,000 tokens, beta footnote', 'Reasoning effort: five levels, default medium, no off switch', 'Price: not published anywhere' and 'Independent score: none yet - no Artificial Analysis row'. The footer reads 'Vendor-reported; nothing independently reproduced. Weights, tech report and model card promised later this month.'

What would change the answer

Beam's case against Gemini rests on the same two items every Beam comparison rests on, and this matchup adds a third.

A price. Without one, the efficiency argument cannot be converted into a budget decision, and the model is competing on a diagram rather than a rate card.

An independent score. Artificial Analysis said on 5 October that Reflection had given it access and that it was benchmarking Beam, describing early indicators as suggesting Beam will be one of the most token-efficient open models it has seen for its level of intelligence. That is the right source saying the right thing, and there is still no Beam row on the board — the model pages return 404 and the leaderboard carries no Beam entry as of this morning.

And the thing that does not change: Beam is text-only. No weight release, no price cut and no index score alters that, which means for a whole class of workloads this comparison will never become a comparison at all. If your pipeline has a video in it, the answer was Gemini 3.1 Pro Preview before the first benchmark loaded.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily