Generated title card titled 'Step 5 Preview vs GPT-6 Astra — ten times the price' with two columns: Step 5 Preview at AA Index 44, 600B total / 27B active, price $1.00 / $2.70, blended $0.51 and 99.8 tokens/sec; GPT-6 Astra at AA Index 53, price $10.00 / $50.00, blended $7.70 and 59.3 tokens/sec. A footer reads 'Both figures per Artificial Analysis Intelligence Index; Astra pricing rises to $20 / $75 above 272K input tokens.'
Guides & Insights

Step 5 Preview vs GPT-6 Astra: Ten Times the Input Price for Nine Index Points

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

These two models were announced seventeen days apart and priced for different planets. Step 5 Preview, StepFun's 600B-parameter sparse mixture-of-experts model announced on September 20, 2026, charges $1.00 per million input tokens and $2.70 per million output tokens. GPT-6 Astra, the vendor's flagship released on September 3, 2026, charges $10.00 and $50.00 for the same units below a 272K-token input threshold — and $20.00 and $75.00 above it. That is ten times the input price and roughly eighteen times the output price for a model that scores nine points higher on the independent Intelligence Index, 53 against 44. Whether Astra is better is not in question. Whether the gap is worth thirty extra dollars per million output tokens on your workload is, and the answer moves twice on the way through this piece.

What each model is for

Step 5 Preview is built and sold for agentic work: AI coding, software engineering, professional knowledge work, finance. The architecture is tuned for it — 600B total parameters with roughly 27B active per token, a 1M-token context, text and image input, and reasoning with extended thinking. StepFun skipped the whole Step 4.x line to get here, moving straight from Step-3.7-Flash to Step 5.

GPT-6 Astra is OpenAI's general-purpose flagship: a 1M-token context, up to 128K output tokens, input in text, images and files, output in text, a knowledge cutoff of April 30, 2026, and support for reasoning, coding and agentic workflows. It is the model an enterprise reaches for when the answer has to be right and the token bill is not the binding constraint.

• Intelligence Index — Step 5 Preview 44 vs GPT-6 Astra 53

• Parameters — Step 5 Preview 600B total / 27B active vs GPT-6 Astra undisclosed

• Context and output ceiling — both 1M-token context; Astra lists a 1,050,000-token window and a 128K output cap, StepFun has published no output ceiling

• Input modality — text and images vs text, images and files

• Output speed — Step 5 Preview 99.8 tokens/sec vs GPT-6 Astra 59.3 tokens/sec

• Price per 1M tokens — $1.00 in / $2.70 out vs $10.00 in / $50.00 out below 272K input, rising to $20.00 / $75.00 above it

• Cache — Step 5 Preview 95% discount vs GPT-6 Astra a 90% read discount with a $12.50 per-million cache-write rate

• Cost per Intelligence Index task — Step 5 Preview $0.71 vs GPT-6 Astra $3.26

Generated two-column comparison scoreboard titled 'Step 5 Preview vs GPT-6 Astra — the scoreboard'. The left column gives Step 5 Preview the rows AA Index 44, parameters 600B total / 27B active, price $1.00 / $2.70, cache discount 95%, cost per index task $0.71 and output speed 99.8 tokens/sec. The right column gives GPT-6 Astra the rows AA Index 53, parameters undisclosed, price $10.00 / $50.00, cache discount 90%, cost per index task $3.26 and output speed 59.3 tokens/sec. A footer reads 'Both figures per Artificial Analysis Intelligence Index; Astra pricing rises to $20 / $75 above 272K input tokens.'

The invoice, line by line

Step 5 Preview's pricing is flat and cheap: $1.00 in, $2.70 out, with a 95% cache discount that puts cached input near $0.05 per million tokens. At a 7:2:1 cache-hit, input and output mix — the convention this blog uses for agent traffic — the blended rate lands near $0.51 per million tokens, and the cost per Intelligence Index task comes in at $0.71, 27th of 200. The caveat is verbosity: the model produced 160M output tokens on that Index against a median of 92M, so raw token counts run above what the ticket price suggests.

GPT-6 Astra's pricing is tiered, and the tier is the part worth reading closely. Below 272K input tokens per request it is $10.00 and $50.00; above it, $20.00 and $75.00. The tier is selected by each request's own input token count, which matters more than it sounds for a model sold on 1M-token context — a single large-context call crosses the line and doubles the rate for the entire request. Cache reads are $1.00 per million, a 90% discount, and cache writes are $12.50 per million. At the same 7:2:1 mix the blended rate lands near $7.70 per million tokens, and the cost per Index task is $3.26, 71st of 200.

Astra cuts the other way on verbosity, and it is the concise model here: 60M output tokens on the Index against Step 5 Preview's 160M, 27th of 200 on that measure. Fewer tokens at a higher rate narrows the gap in a way the raw multiple overstates. It does not close it — the cost-per-task figure already nets out verbosity, cache behaviour and pricing together, and $3.26 against $0.71 is still a 4.6-times difference.

What the nine index points buy

Astra's headline numbers are OpenAI's own and unreproduced: 96.1 on GPQA Diamond, 72.6% on OSWorld 2.0, 97.6% on FrontierMath Tier 4, 57.2% on Humanity's Last Exam with tools, and around 57.9% on Terminal-Bench 4.0. The ARC-AGI-3 figure is the one to handle carefully — OpenAI reports 99.9% under its own provider-adapter harness and 62.7% on the standard harness, and a 37-point swing between harnesses is a statement about the harness at least as much as about the model.

Step 5 Preview's published numbers sit in the same territory on the agentic tasks it was built for, and carry the same caveat: 33.3% on Terminal-Bench 4.0, 80.5 on ProgramBench, 67.7 on DeepSWE v1.1 and 49.0 on StepCodeBench, all from StepFun and none reproduced independently. Different vendors, different harnesses, no shared run — which is precisely why the independently measured nine-point gap is the more useful number than any of these.

That gap is real and it is not small. Across a 200-model board, 53 and 44 sit at third and twenty-fourth position, and the separation shows up as a different class of task rather than a different score on the same task: Astra's published wins cluster on long-horizon reasoning and computer use, which is where the top of the board pulls away. If your workload lives there, the nine points are worth more than the invoice.

One caveat on Astra's score. Artificial Analysis revises the Intelligence Index, and figures for the same model drift across snapshots taken weeks apart — 52.8, 53 and 54.7 all appear in material published about this model within the same month. The 53 on the current model page is the number to use, and any comparison that sets an older Astra snapshot beside a current Step 5 Preview figure is measuring with two different rulers.

Screenshot of the Artificial Analysis model page for GPT-6 Astra, showing OpenAI as the creator with a September 2026 release date, an Intelligence Index of 53 at rank 3 of 200, output speed 59.3 tokens per second at rank 101 of 200, cost per Intelligence Index task $3.26 at rank 71 of 200, verbosity 60M output tokens at rank 27 of 200, pricing of $10.00 per million input tokens and $50.00 per million output tokens with a 90% cache discount, and technical specifications listing reasoning, text and image input, a 1M-token context window and an April 30, 2026 knowledge cutoff.

Speed and latency are two different questions

On generation speed the independent board is unambiguous: 99.8 output tokens per second for Step 5 Preview against 59.3 for GPT-6 Astra, which is also slower than the 70 tokens-per-second average across the models measured. In an interactive coding loop that is the difference between a suggestion and a pause, and it compounds across a session the way a cache discount does.

Latency is the more complicated story, and a single number misleads on it. Astra routes reasoning before it emits, so time to first token and time to a usable answer are not the same measurement, and published figures for it spread widely depending on whether reasoning is counted. On our own traffic over the last seven days, GPT-6 Astra's median time to first token is 6.72 seconds with a 95th percentile of 10.00 seconds — the number a caller actually experiences through our endpoint. Artificial Analysis puts Step 5 Preview's time to first token at 2.96 seconds on its own harness, and those two figures are not measured the same way, so they should not be subtracted from each other.

Where the cache asymmetry actually lands

Both models discount repeated prefixes heavily and both charge for writing them, and the difference between them is 20-fold on the read. Step 5 Preview's 95% cache discount puts cached input near $0.05 per million tokens. GPT-6 Astra reads cache at $1.00 per million — a 90% discount off a base rate ten times higher — and writes it at $12.50 per million.

For an agent loop re-sending the same system prompt and repository context on every turn, the read rate is what compounds, and the deeper discount on the cheaper model is worth more than the shallower discount on the expensive one. Astra's blend still lands near $7.70 per million tokens against Step 5 Preview's $0.51 at the same 7:2:1 mix — a fifteen-fold difference that a long session does not close but widens.

Running both from one key

GPT-6 Astra is live on OrcaRouter under openai/gpt-6-astra at OpenAI's list rates — $10.00 and $50.00 per million below the 272K threshold, $20.00 and $75.00 above it — passed through with zero markup, so a vendor price change is live on our side the same day rather than at the next billing cycle. Our own seven-day telemetry on that endpoint shows the 6.72-second median and 10.00-second 95th-percentile time to first token quoted above, alongside 277.9M tokens of traffic through it in the last week.

Step 5 Preview is not in our catalogue and we do not route it. It runs on StepFun's own API and Studio, and that is the only place it runs until the October 15 checkpoint lands. That asymmetry is not a flaw in the comparison; it is the comparison. The cheap half of this matchup is the half you have to go elsewhere to call, and the half you can put into production today sits behind one API for more than 200 models: GPT-6 Astra at the rates above, GLM-5.3 at $1.26 and $3.96, DeepSeek V4 Pro, Claude Opus 5 and the rest, with automatic failover across providers holding a path steady when a provider's capacity moves and the routing DSL letting a task start cheap and escalate only when it has to. That escalation rule is the real answer to a ten-times price gap: most of a workload does not need the top of the board, and a routing layer is how you stop paying for the part that does not.

Screenshot of the OrcaRouter model page for GPT-6 Astra at www.orcarouter.ai/models/openai/gpt-6-astra, showing the model id openai/gpt-6-astra with NEW, FLAGSHIP and FEATURED badges, a 1M-token context and 128K max output, text, image and file input with text output, best-for tags for reasoning, coding and agentic work, input pricing of $10.00 and output pricing of $50.00 per million tokens, a seven-day median time to first token of 6.72 seconds and a 95th percentile of 10.00 seconds, 277.9M tokens of seven-day traffic, and the chat-completions and responses endpoints behind the OrcaRouter API.

Who should pick which

If your workload is a long-horizon agent that has to get a hard task right — computer use, multi-step research, work where a wrong answer costs more than the tokens — GPT-6 Astra's nine-point lead is the cheapest part of the decision, and the invoice is what the capability costs. Run it as the escalation target, not the default, and watch the 272K threshold: a single oversized request doubles the rate for everything in it.

If your workload is high-volume agentic text — code generation, refactors, structured extraction, the inner loop of a coding assistant — Step 5 Preview is ahead on everything the invoice and the clock care about, and nine index points are unlikely to survive contact with a task distribution that does not sit at the top of the board. The catch is not performance: the model is proprietary, with no published licence, no commercial-use grant and no checkpoint until October 15. You can call it; you cannot own it, fine-tune it or ship a derivative of it.

If the answer is not obvious, the pattern is a tier split rather than a choice. Route general traffic to a mid-tier model and reserve Astra for the requests that need the top of the board — both ends of that rule live behind one key on our side, so the split costs a routing rule rather than a second contract. Paying flagship rates for work a mid-tier model finishes correctly is the mistake this comparison is actually about.