Generated title card titled 'Step 5 Preview vs Gemini 3.1 Pro — two previews, seven months apart', with a timeline marking Gemini 3.1 Pro Preview at February 19, 2026, still preview, and Step 5 Preview at September 20, 2026, weights due October 15. Below, two cards: Step 5 Preview with AA Index 44, input text and image, and time to first token 2.96s; Gemini 3.1 Pro Preview with AA Index 30, input text, image, audio, video and file, and time to first token 35.72s. A footer reads 'Both per Artificial Analysis Intelligence Index v4.3.2.'
Guides & Insights

Step 5 Preview vs Gemini 3.1 Pro: Two Models Called Preview, Seven Months Apart

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Both Step 5 Preview and Gemini 3.1 Pro carry the word preview, and the word does not mean the same thing on either. StepFun announced Step 5 Preview on September 20, 2026, with the API open, the weights scheduled for October 15, and a BF16 repository on Hugging Face containing nothing but a .gitattributes. The other has been shipping as gemini-3.1-pro-preview in the API and as "Gemini 3.1 Pro Preview" on every leaderboard since February 19, 2026, handling production traffic for seven months without the label ever coming off. One is a genuine preview of something not finished. The other is a naming convention attached to a finished product. Comparing them on the same axis means deciding which of those two things you are actually buying, and the second-order answer — that the model still called a preview after seven months is the more predictable of the two — is the part worth carrying into a procurement decision.

What the same board says

Artificial Analysis scores both on Intelligence Index v4.3.2, ten evaluations. It is the only place these two have been measured by the same hand, and the result is further apart than the vendor materials on either side would suggest.

• Intelligence Index — Step 5 Preview 44 vs Gemini 3.1 Pro Preview 30

• Rank — Step 5 Preview 24th of 200 models vs Gemini 3.1 Pro Preview 69th of 200

• Price per 1M tokens — Step 5 Preview $1.00 in / $2.70 out vs Gemini 3.1 Pro $2.00 in / $12.00 out

• Cache discount — Step 5 Preview 95% vs Gemini 3.1 Pro 90%

• Output speed — Step 5 Preview 99.8 tokens/sec vs Gemini 3.1 Pro 118.9 tokens/sec

• Time to first token — Step 5 Preview 2.96s vs Gemini 3.1 Pro 35.72s

• Context — both 1M tokens; our catalogue lists Gemini 3.1 Pro with a 65K output ceiling, StepFun has not published one

• Input modalities — Step 5 Preview: text and image. Gemini 3.1 Pro: text, image, audio, video and file. Both output text only

• Weights — Step 5 Preview closed today, BF16 checkpoint scheduled for October 15; Gemini 3.1 Pro proprietary throughout

A 14-point gap on a scale where the top of the board sits at 53 is a large one, and it should be reported next to the thing that makes it strange: Google's own published evaluation numbers for this model are excellent. The vendor reports 77.1% on ARC-AGI-2, 82.8% on its multimodal suite, 94.1 on GPQA Diamond and 47.0 on Humanity's Last Exam. Those are not weak figures. They are also Google's own, run on Google's harnesses, and they measure specific capabilities rather than the aggregate the index scores. Both sets of numbers can be accurate at once. When a vendor's headline evaluation and an independent composite disagree this sharply, the composite is the one to plan a production workload against, because it is the one that penalises the model for the things the vendor did not choose to measure.

The two speed rows are worth reading together rather than separately. Gemini 3.1 Pro generates tokens 19% faster once it starts — 118.9 against 99.8 — and takes 35.72 seconds to start, against Step 5 Preview's 2.96. That is the profile of a model that deliberates before it emits. For a batch job where a human is not waiting, the faster generator wins on throughput. For an agent loop making dozens of dependent calls, a twelve-fold difference in time to first token is the difference between an interactive tool and a queue.

Screenshot of the Artificial Analysis model page for Step 5 Preview, showing StepFun as the creator and a September 2026 release date, an Intelligence Index of 44 at rank 24 of 200, output speed 99.8 tokens per second, cost per Intelligence Index task $0.71, verbosity 160M output tokens, pricing of $1.00 per million input tokens and $2.70 per million output tokens with a 95% cache discount, and technical specifications listing reasoning, text and image input, and a 1M-token context window.

The one axis where Gemini wins outright

Modality is not a footnote in this comparison. Gemini 3.1 Pro accepts text, images, audio, video and arbitrary files, and Step 5 Preview accepts text and images. If your pipeline involves a screen recording, a call recording, a PDF bundle or a video feed, only one of these two models can take the input at all, and the 14-point index gap is irrelevant because the other model cannot answer the question.

That asymmetry decides more real workloads than the benchmark tables do. Video review, meeting analysis, audio transcription plus reasoning, and document workflows built on scanned files are all cases where Gemini 3.1 Pro's breadth makes it the only candidate on this page. It is also why the model has stayed in production for seven months under a preview label: the workloads it holds are the ones where a newer text-and-image model cannot displace it.

For text and image work, the calculus inverts. Step 5 Preview is 14 points ahead on the aggregate, roughly twice as fast on input price and 4.4 times cheaper on output, and answers in a twelfth of the time. On a document-extraction or chat-over-screenshots workload, there is no reason to pay the premium and wait the extra eleven seconds of deliberation.

Where the pricing gets less flat than it looks

The headline prices understate the difference, and the reason is a tier boundary rather than a rate.

Gemini 3.1 Pro's $2.00 and $12.00 hold up to 200,000 input tokens. Above that, both rates step to $4.00 and $18.00 — a 50% increase on output that applies to the whole request, not just the portion past the boundary. Step 5 Preview's $1.00 and $2.70 have no published tier; the price is the price at any length the context window allows.

Both models advertise a 1M-token context, so both invite you to build long-context pipelines. On one of them, a 300,000-token request costs twice what the price list suggests; on the other it does not. That is a structural advantage for Step 5 Preview in exactly the category StepFun built it for — long-horizon agentic work with a large, repeatedly-referenced context — and it is compounded by the cache discount, where Step 5 Preview's 95% against Gemini 3.1 Pro's 90% widens the gap further on the repeated-prefix pattern an agent loop produces.

Worked roughly, and flagged as arithmetic rather than a quoted figure: an agent loop that sends a 250,000-token context on every one of forty turns is ten million input tokens. At Gemini 3.1 Pro's above-200K rate, with the cache absorbing the repeated prefix, that is a meaningful step up from what the $2.00 list rate implies. At Step 5 Preview's flat $1.00 with a deeper cache discount, the same loop costs less and, more importantly, costs something you can predict from the price list without reading the tier table first.

Generated two-column comparison scoreboard titled 'Step 5 Preview vs Gemini 3.1 Pro — the scoreboard'. The left column gives Step 5 Preview the rows AA Index 44, price $1.00 / $2.70, cache discount 95%, time to first token 2.96s, context 1M tokens with no published output cap, and input text and image. The right column gives Gemini 3.1 Pro Preview the rows AA Index 30, price $2.00 / $12.00, $4.00 / $18.00 above 200K input, cache discount 90%, time to first token 35.72s, context 1M tokens with a max output of 65K, and input text, image, audio, video and file. A footer reads 'Both per Artificial Analysis Intelligence Index v4.3.2; Gemini ARC-AGI-2 and multimodal figures are vendor-reported.'

Availability is the quiet difference

Gemini 3.1 Pro has been callable for seven months through Google's API and, more usefully for planning, through multiple independent providers — which is what makes a p50 latency figure meaningful rather than anecdotal. Step 5 Preview opened its API on September 20 and has exactly one source. A single-source endpoint that is days old is the least predictable component you could put on a production path, not because the model is bad but because nobody knows its capacity curve yet.

That is the argument for routing rather than choosing. Gemini 3.1 Pro Preview is in our catalogue at $2.00 and $12.00 with the tier step passed through unchanged, and the models it is measured against sit beside it, behind one key for more than 200 models with provider list price passed through at 0% markup. Step 5 Preview is not on that list — it runs on StepFun's own API, and until the weights land that is the only place it runs. Automatic failover is what makes a days-old preview safe to trial rather than a bet: keep the production path on a model with a track record, send the text-and-image bulk to Step 5 Preview while you evaluate it on your own traffic, and let the fallback hold when the preview endpoint degrades. There is no second contract and no separate SDK for the models we do route, so a Google repricing shows up on your bill as the current number rather than at renewal.

Screenshot of the OrcaRouter model page for Gemini 3.1 Pro Preview at www.orcarouter.ai/models/google/gemini-3.1-pro-preview, showing the model id google/gemini-3.1-pro-preview with Vision, Audio, Tools, JSON and Reasoning tags, a 1M-token context, a 65K max output and text output, input pricing of $2.00 and output pricing of $12.00 per million tokens, a p50 time to first token of 10.00 seconds, 109.7M tokens of traffic over seven days, and the endpoints and SDK snippets behind the OrcaRouter API.

Which one you are actually buying

If your work arrives as video, audio or files, Gemini 3.1 Pro is the only model on this page that can accept it, and the index gap does not enter the decision. Seven months in production with the preview label still attached is a stability signal, not a warning: the naming convention persists because Google has never needed to change it, and the model behaves as the production workhorse it is.

If your work is text and images at volume with a large repeated context, Step 5 Preview is ahead on every metric that matters — 14 index points, half the input price, a 4.4x cheaper output rate, no tier cliff, a deeper cache discount, and a time to first token measured in seconds rather than in half a minute. What you are buying there is a days-old single-source endpoint with no published licence and no published output ceiling, which is a real risk and a bounded one.

The thing to watch is not the index. It is whether StepFun publishes an output ceiling alongside the 1M-token context window, because the vendor's announcement is silent on it and a separate third-party configuration that circulated during the leak listed 350,000 context with a 64,000 output cap — a materially different product from the one on the price list. Gemini 3.1 Pro's 65K ceiling, as our catalogue lists it, is not generous either, but it is stated, and a stated limit is one you can design around.