OpenAI Astra-1
Guides & Insights

OpenAI Astra: Everything We Know About the Model That Isn't GPT-6

Author

Jim Song

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The thing you can download today is not a model. It is ten files of Lean code. On 1 August 2026, OpenAI published a research post about mathematics, and buried in it was the first confirmation of the name of its next frontier system: Astra. There is no model card, no pricing page, no API endpoint and no date. The current flagship you can actually call is still GPT-5.6 Sol, at $5 per million input tokens and $30 per million output tokens, exactly as it has been since 9 July.

If you searched for GPT-6, here is the short version: OpenAI has never announced a model by that name, and as of today it has not decided whether Astra will ship as GPT-6, as a GPT-5 point release like GPT-5.7, or as a fourth tier sitting alongside Sol, Terra and Luna. That naming ambiguity is not press-shyness. It is a live internal question, reported by The Information on 31 July and not resolved since.

What follows separates three things that most coverage of this story runs together: what OpenAI itself stated, what credible reporting adds, and what is circulating with no source attached. The proofs are real artifacts you can compile. The 10-trillion-parameter figure is a number someone typed on the internet. Those deserve very different weight.

What OpenAI actually said — and the much longer list of what it didn't

The announcement did not arrive as a product launch. It arrived as a paper. OpenAI's post described results produced by "an internal version of Astra, our next major model," across problems that had, in its framing, seen no progress on the main result for at least a decade. Astra was characterized as a system built for long-running work: multiple agents coordinating on different parts of one large problem over extended stretches rather than answering in a single pass.

That is close to the entire official technical description. Set against it, the list of absences is striking:

Confirmed by OpenAI — the name Astra; that it is the "next major model"; that an internal version produced ten results; the multi-agent, long-horizon design intent; a 249-page manuscript; Lean 4 certificates on a public repository; a token-cost estimate quoted at GPT-5.6 Sol rates.

Not stated by OpenAI — release date; price; context window; parameter count; any benchmark score; the model card; whether it reaches ChatGPT or only the API; how many agents ran, for how long, on what hardware; the prompts; the success rate; the final product name.

Credible reporting fills in a little. The Information reported that Sam Altman demonstrated Astra to policymakers in Washington in late July — a briefing that, per that reporting, included Senators Raphael Warnock, Bernie Moreno and Mark Warner, with Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick also present. The model family is described as already in testing, with "Astra" itself still a tentative label.

Nobody outside OpenAI has run this model on anything. Every capability claim in circulation traces back either to the ten published proofs or to a demo the public did not see.

The ten proofs: unusually checkable, and still not settled

OpenAI Astra-2

This is where the announcement is genuinely stronger than a typical benchmark drop, and it is worth being precise about why. OpenAI did not publish a score and ask to be believed. It published machine-checkable artifacts under Apache 2.0 in the openai/ten-proofs repository — one Lean file per result, pinned to Lean 4.32.0 with a mathlib dependency, alongside a ComparatorChallenges directory for independent checking. The repository landed as a single commit on 1 August and has since picked up several hundred stars and dozens of forks.

The ten results, as the repository names them, span an unusually wide spread of fields:

Group theory — the construction of a non-sofic group, answering a question open since 1999 in the negative. Soficity is a property nearly every group mathematicians work with satisfies; whether every countable discrete group must satisfy it had stood for 27 years.

Operator algebras — a counterexample bearing on the Connes rigidity conjecture, posed by Fields medalist Alain Connes in 1980.

High-dimensional geometry — an improvement to the sphere-packing method, reported as the first of its kind since 1978 — plus a sharp form of Ehrhart's volume inequality.

Coding theory — new bounds for binary and spherical codes.

Complexity — an arithmetic circuit lower bound for the permanent, and an exponential parallel-repetition result for entangled games.

Lattice cryptography — fixed-polynomial hardness for the closest vector problem.

Extremal combinatorics — the growth scale of multicolor triangle Ramsey numbers, and counterexamples in extremal graph theory, including work on several catalogued Erdős problems.

Reactions from mathematicians were interested rather than dismissive. Thomas Bloom, who maintains the Erdős problems database, called the results big news and rated them above the unit-distance conjecture counterexample from May. OpenAI's Noam Brown offered the useful deflation from the inside: "Sadly, no Millennium Prize Problems (yet)."

What a Lean certificate settles, and what it leaves open

A Lean proof that type-checks is valid by construction. You do not have to trust OpenAI's evaluation methodology, its choice of baselines, or its interpretation of its own results — you can compile the files. For an industry where "we scored 94.1%" is the standard unit of evidence, that is a meaningful upgrade in falsifiability.

It is also narrower than the headlines imply, in four specific ways. Lean checks that the proof of a formal statement is sound. It does not check that the formal statement is the theorem the prose claims. It does not check that the encoded definitions match what the field means by those words. It does not check the informal reductions bridging the formal endpoints to the headline result. And it says nothing about novelty — whether a result is new, or already known under a different name.

All four of those are human-judgment questions, and as of the announcement none had gone through refereed review. The manuscript acknowledges consultations with specialists; acknowledgments are not endorsements of every claim. One published partial build of the repository compiled 8,820 of 9,007 jobs, which is what a large real formalization looks like mid-verification rather than a finished audit. The honest status today is ten serious, formally encoded research claims — not ten settled entries in the mathematical canon.

There is a second, quieter limit: the discovery process is not reproducible by anyone outside OpenAI. The proofs are public; the search system, prompts, checkpoints and execution environment are not. You can verify the destination. You cannot rerun the journey.

The $2,000 figure, and why it buys less than it sounds like

OpenAI Astra-3

The number that traveled furthest was the cost. OpenAI put the token spend behind the ten solutions at roughly $2,000 priced at GPT-5.6 Sol rates — about $200 per decade-old open problem, on the most common reading of its phrasing. Some coverage read the same sentence as under $2,000 per problem, a tenfold difference; OpenAI's post is terse enough that both readings survived into print, and we have not seen the company disambiguate it.

Take the total reading and work it out. Sol bills $5 per million input tokens and $30 per million output, with reasoning tokens billed as output. If the spend were entirely output, $2,000 buys at most about 67 million output tokens — roughly 6.7 million per problem. That is a lot of thinking, and it is not an absurd amount of thinking. It is the sort of budget a well-funded research group could already spend on a single hard question.

Three caveats matter more than the headline:

It is a counterfactual price, not a price. Astra is unreleased and unpriced. The $2,000 describes what those tokens would cost at another model's rates. A frontier system built for hours-long multi-agent runs has no obvious reason to be priced like Sol, and every reason to be priced above it.

It counts only the wins. No success rate was published. We do not know how many problems Astra was pointed at and failed to crack, or what those attempts cost. A cost-per-success published without a success rate is not a cost-per-result, and the difference could be an order of magnitude in either direction.

The tokens are the cheap part. Humans framed the problems, prepared the manuscripts and drove the formalization. That labor is not in the $2,000.

The one part of this you can price today is the baseline. Because OrcaRouter passes provider list price straight through at 0% markup, GPT-5.6 Sol costs the same $5/$30 on our API as it does on OpenAI's — so if you want to sanity-check the arithmetic against your own workload, the rate in the announcement is the rate you would actually pay. What no one can price yet is Astra itself, and any vendor telling you otherwise is guessing.

The release date is gated by a 30-day clock, not by a leak

OpenAI Astra-4

Most "when is GPT-6" coverage is rumor arithmetic. There is a more concrete constraint sitting in plain sight, and it has barely been connected to the question.

An executive order signed on 2 June 2026 directed federal agencies to stand up a voluntary review process under which frontier developers submit models for government review up to 30 days before public release. The framework was slated to take formal effect on 1 August — the same day the Astra post went up. Reporting indicates Astra is expected to be the first model through it.

"Voluntary" is doing some work in that sentence. In practice the administration has leaned on release timing before, and OpenAI's own recent behavior looks like a rehearsal: GPT-5.6 was held to roughly twenty vetted organizations from 26 June before broad access arrived on 9 July — a staged rollout of about two weeks.

Stack those two facts and you get a rough floor rather than a prediction. If Astra goes through a 30-day pre-release review and then repeats the GPT-5.6 pattern, broad availability lands something like six weeks after submission. And the submission date is precisely what we do not know. This is why confident August dates should be discounted: the gating step is a process with a published duration, and the clock has not visibly started.

The prediction markets have drifted in exactly that direction. Read on 5 August 2026, Polymarket's "GPT-6 released by" ladder sat at 1% for 7 August, 3% for 14 August, 13% for 21 August, 25% for 31 August, 68% for 30 September and 89% for 31 December, on close to a million dollars of volume. Treat that as a crowd's aggregated guess, not a source — but note that the crowd has moved its weight to Q4.

It also helps to remember the track record. Every "GPT-6 launches next week" claim of the past twelve months has been wrong. An unverified leak pegged 14 April 2026 as launch day; what actually shipped, nine days later, was GPT-5.5.

The rumors, graded

Sorted by how much weight each actually carries:

Naming undecided (GPT-6 / GPT-5.7 / a separate class). Sourced to The Information. The most reliable non-OpenAI claim in this story, and the one that should reset expectations — a system announced through a math paper may well not be branded as a generational jump at all.

Codenames "Zinc" and "Magnesium" spotted in Design Arena. Community sighting, entries reportedly disabled, unconfirmed by OpenAI. Read as: a GPT-5.x point release may well land before Astra, which would satisfy nobody's GPT-6 expectations while absorbing the news cycle.

Roughly 2× GPT-5.6 Sol in scale; pricing near the GPT-5.6 range. Circulating rumor. The leaker behind it acknowledged not having tested the model. Plausible-sounding and entirely unsubstantiated.

A 1.5-million-token context window. Appears in leak roundups with no named source. Worth noting Sol already ships 1,050,000 tokens, so this would be an increment, not a leap — which is a reason to be suspicious of it as a headline "leak."

Approximately 10 trillion parameters. No attribution anywhere we could trace. The most-repeated and least-supported number in the entire story.

Memory and personalization as the defining direction. Grounded in Altman's own public remarks from an August 2025 interview — real, but a year old and about the post-GPT-5 direction generally, not about Astra specifically.

What to actually do between now and whenever it ships

The wrong move is rearchitecting for a model that has no card. The right move is noticing which of your problems a long-horizon, multi-agent system would change, because those are the parts of your stack that are underbuilt today regardless of what OpenAI ships.

Three of them are worth building against GPT-5.6 Sol right now:

Cost ceilings per task, not per call. If a single job can run for hours and spend six or seven million output tokens, per-request limits stop protecting you. You need a budget that a whole task inherits, and a hard stop.

Checkpointing and resumability. A one-shot call either returns or fails. An hours-long agent run that dies at minute 90 with nothing durable written is a category of expensive failure most codebases have never had to handle.

Evaluation you trust on problems with no reference answer. The ten proofs are interesting partly because Lean supplies an oracle. Almost nothing in production does. If you cannot tell a good six-hour run from a plausible-looking bad one, more compute will not help you.

On the switching question: Astra is not available on OrcaRouter, because it is not available anywhere. What we can say is that the version-churn problem is largely solved on our side. One key reaches 200+ models, so whether this ships as GPT-6, GPT-5.7 or a fourth GPT-5.6 tier, adopting it is a model-string change rather than a migration — no second contract, no new SDK. And for an unproven frontier model specifically, automatic failover is the difference between evaluating it in a real workload and betting a production path on a system nobody outside the lab has stress-tested. You route a slice of traffic to it, keep Sol underneath as the fallback, and find out.

Questions worth answering

Is Astra the same thing as GPT-6?

Not necessarily, and that is the most consequential unknown in this story. OpenAI has confirmed Astra as its next major model but has not decided its release name; reporting places GPT-6, GPT-5.7 and a separate class alongside Sol, Terra and Luna all on the table. It is entirely possible that a model named GPT-6 never appears, and that the capability jump people have been waiting for arrives under a name nobody was tracking.

Can I access Astra in any form today?

No — not in ChatGPT, not through the API, not through any router or reseller. What exists publicly is the paper, the reasoning walkthroughs and the Lean repository. If a service claims to offer Astra access, it is either reselling something else or lying. The only genuinely useful thing you can do with the release right now is compile the proofs yourself.

Did Astra really solve problems human mathematicians couldn't?

It produced formally verified proofs of statements that had been open for at least a decade, which is a real and unusual result — and the verification is a step above the industry norm. But "Lean-checked" is not "peer-reviewed," and the framing question of whether each formal statement captures the headline problem is exactly the part Lean cannot answer. Specialists reading the manuscripts have been positive; none of it has been through refereed review. Expect the assessment to firm up over months, in either direction.

Does the $2,000 figure mean frontier research is now cheap?

It means the winning token runs were cheap, at another model's prices, excluding the failures and excluding the humans. Each of those three exclusions could be large. The honest read is that the marginal cost of a successful automated proof search has fallen far enough to be interesting, not that ten open problems now cost the price of a laptop.

What would actually settle this

Three specific things, none of which are rumors, and all of which are checkable when they happen.

A model card and a pricing page. That is the moment Astra stops being a research artifact and becomes something you can plan against — and the moment the naming question resolves itself.

An independent rebuild of the Lean repository by specialists who audit the definitions and the informal reductions, not just the type-checking. The ComparatorChallenges directory suggests OpenAI expects this. If the formal statements hold up under that scrutiny, the results get much harder to discount; if a mismatch surfaces in even one of the ten, every claim in the set gets re-read.

Evidence that the federal review clock has started. Under a 30-day pre-release window, that submission is the earliest reliable signal of a ship date — considerably more informative than the next screenshot of a disabled entry in an arena leaderboard.

Until then, the accurate statement is narrow and worth holding onto: OpenAI's next flagship has a name, ten machine-checkable proofs, a Washington demo, and nothing else. Anyone offering you a date is guessing, and anyone offering you a parameter count is making it up.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube