
OpenAI's Aeon Agent Leak: What the Rumoured Grok Bot Rival Actually Is, and What Is Still Unverified
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
Something called Aeon is reportedly about to be released by the vendor — a persistent, always-on agent positioned against Grok Bot, the digital-coworker product SpaceXAI shipped on August 11, and against Meta's Muse. The claim is circulating in a form that is worth reading carefully: the vendor's "Grok Bot" ("Aeon") is nearing release, possibly before DevDay, possibly this week. There is no product page, no pricing, no documentation, no benchmark, and no public statement naming the word. What does exist is thinner and more interesting than the summary suggests: two model-ID strings spotted inside a Codex Desktop feature flag on September 3 — "gpt-6-astra" and "gpt-6-astra-aeon" — sitting next to a confirmed platform launch, a confirmed conference date of September 29, and a confirmed flagship model, GPT-6 Astra, that shipped September 3 and took the top of the company's line from GPT-5.6 Sol. Aeon is the most plausible next thing on the vendor's roadmap. It is also, as of today, a string in a config file and a set of reports that all trace back to the same few posts.
What the leak actually consists of
It is worth separating the artifact from the reporting, because the two are not the same weight.
The artifact is concrete and dated. On September 3, a screenshot circulated showing two undocumented model identifiers added to a Statsig feature flag inside Codex Desktop: "gpt-6-astra" and "gpt-6-astra-aeon". The image came from the account @notjazii and was amplified by @kimmonismus — the same account whose post is the signal behind this article — with the observation that the models were "already" in Codex and could appear "any minute now." The same screenshot is what first tied OpenAI's confirmed Astra model to the GPT-6 name. That is the whole of the primary evidence, and it is genuine: feature flags are how vendors stage configurations before a launch, so the strings are real deployment preparation rather than invention.
What they do not do is establish that anything is available, or what it is. A staged identifier can sit unused, get renamed, or refer to an internal experiment that never ships under that name. Nothing in the screenshot describes behaviour, memory, context limits, price, or who gets access. The suffix itself is ambiguous: it could be a longer-running agent configuration, a product tier, a safety profile, or a flag for an internal evaluation build. Treating "gpt-6-astra-aeon" as a public model called Aeon runs ahead of the evidence, and that is the honest reading of the artifact even though the artifact is real.
The reporting is where the detail comes from, and it is uniformly secondary. Reports describe Aeon as a 24/7 agent that retains memory across days, advances long-horizon projects in the background, resumes on event triggers, and runs on the Astra model line. Leaked interface fragments are said to show memory, scheduled actions, and third-party app connectors, with commerce flows — product research, cart building, post-purchase follow-ups — inside one thread. Claims of sub-second tool calls and native support for Indian payment and logistics APIs appear in a single report with no corroboration and no screenshot attached. A separate thread of reporting links the name to OpenAI's internal math work: reports through September say an internal model, speculated to be Aeon, was used in long-horizon runs that produced progress on Millennium Prize problems, with one account describing roughly 10,000 coordinating agents and an 88-hour run on Navier–Stokes with Lean formal verification. OpenAI has confirmed progress on a second Millennium problem; it has not confirmed that Aeon, or any named model, was behind it. Those are two different facts and the reporting regularly merges them.
One more piece of the summary deserves a correction, because it recurs. The same signal text mentions GPT-6 Sol as expected imminently, with Claude Opus 5.5 and Claude Sonnet 5.5 as Anthropic's response. Sol is a leaked name with one third-party artifact behind it — the string "GPT-6 Sol medium" in an NVIDIA merge record — and it is not OpenAI's shipping flagship; GPT-6 Astra is, since September 3, alongside Astra Pro. The Claude 5.5 names are unconfirmed strings that appear in prediction-market qualifying lists. If you are reading a launch calendar assembled from those, it is a calendar of rumours, not of ships.
What OpenAI has actually shipped, and what that says about Aeon
The strongest evidence for Aeon is not the leak at all — it is the two weeks of OpenAI releases sitting on either side of it.
• September 3 — GPT-6 Astra reached general availability, OpenAI's frontier model, 1M-token context, listed at $10.00 input and $50.00 output per million tokens with a 90% cache discount. (Artificial Analysis dates the release to September 3; our own catalogue lists the model under 2026-09-04.)
• September 10 — OpenAI opened the Agents API in public beta, exposing the managed Codex harness: OpenAI runs the session orchestration, context compaction and recovery, while developers supply the task, model, tools and execution environment, with sandboxes hosted by OpenAI, self-hosted, or on partner infrastructure. Billing is tokens, tools and container time, with no separate platform fee.
• September 10 — GPT-Live-1 arrived in the API at $0.05 per voice minute.
• September 15–16 — Sam Altman posted that OpenAI had a "big ship this week, and then for devday ship x 6", then walked the first half back: the thing he was most excited about launching that week would land the following week instead.
• September 29 — DevDay 2026 at Fort Mason Center, San Francisco, confirmed since June, keynote livestreamed.
Read those together and Aeon stops looking like a wild rumour. OpenAI has just productised the agent harness and sold it to developers; the consumer-facing version of that same capability, sold as a subscription against Grok Bot and Muse, is the obvious next move, and Altman's "next week instead" lands squarely on the window between September 16 and DevDay. None of that makes the Aeon name verified. It does mean the shape of the story is coherent: an agent product, on the Astra line, around DevDay.

If Aeon ships, what is it actually competing with?
Three races, not one — and they have different scoreboards.
The persistent-agent race is already lost on time. Grok Bot went into early beta on August 11 at a $120-per-seat entry point, bundled into SpaceXAI's Cursor Premium Teams tier, with individual tiers at $200 and $300 a month. Each bot gets its own cloud VM with a browser, filesystem and terminal, signs into apps that have no API, saves demonstrated workflows as routines, and keeps working after you close the laptop. Meta's Muse followed on September 8. OpenAI would be third to a category it defined the research for.
The capability race is genuinely open, because nobody has published a number. SpaceXAI has released no independent benchmarks or reliability data for Grok Bot; the model router behind it is automatic and users cannot pin a model; testers have been publicly unimpressed; and reviewers note there is no enterprise control layer — no scoped permissions, escalation rules or audit trail, with credential isolation described as per-user rather than per-bot. That is a low bar, and it is the same bar Aeon would clear or fail on, invisibly, because agent products are not benchmarked the way models are. There is no Artificial Analysis index for "did the agent finish the task without doing something catastrophic."
The pricing race is the one readers will feel. Grok Bot's $120 is a floor, not a ceiling — usage beyond the included limits bills at token cost. If Aeon arrives inside an existing OpenAI subscription tier, the comparison stops being about list price and becomes about which seat you already pay for. If it arrives as its own tier, OpenAI has to justify a new line item against a product that has been shipping for six weeks.
And there is a fourth constraint that the leaks never mention and the reporting should. Astra reached OpenAI's Critical cybersecurity capability threshold, and the company has warned that monitoring may slow, pause or stop legitimate tasks — explicitly including long-running jobs. A persistent agent is a long-running job by definition. Whatever Aeon is, it ships into a safety posture that OpenAI has said can interrupt exactly the workloads the product is sold on.
The dated tests that would settle this
Rumours about a product a week out are cheap. These are the specific things that would turn Aeon from a string into a product, and each is checkable:
• A public OpenAI page or release note that uses the name Aeon, rather than a screenshot of a config file. This is the single cleanest signal and it has not appeared.
• An entry in OpenAI's own model or product catalogue, or a model identifier that resolves through the API. A Statsig flag is not that.
• Any published price, seat tier, or access rule. The reporting currently says Plus and Business first with API access trailing; that has not been sourced to OpenAI.
• The DevDay keynote on September 29. This is the natural reveal surface, and if nothing appears there the "before DevDay" framing collapses on its own.
• A capacity signal. OpenAI paused new $200 Pro subscriptions in early September to protect Astra capacity under demand; a new always-on agent product consumes inference continuously rather than per session, so availability limits are a real question the rumours do not touch.
How to hedge a product that does not exist yet
There is no Aeon endpoint to call, so there is nothing to switch to — and this is the point in the cycle where the useful move is not "adopt" but "make the switch cheap for later."
Two things you can do today that cost nothing and pay off whichever way this lands. First, keep the model layer behind one endpoint rather than scattered across per-vendor SDKs: OrcaRouter puts the OpenAI lineup — GPT-6 Astra, GPT-5.6 Sol and the rest of the GPT-5.6 line — and 200-plus models from other vendors behind a single API key at provider list price with 0% markup, so a new model or a new agent surface is a config change rather than an integration project. Second, if you are going to trial an agent product the week it launches, trial it somewhere a failure is survivable: automatic failover across providers means an unproven route can sit next to a model you have already characterised instead of replacing it, and the routing DSL lets you compose several models into one call so an agent loop can use a cheap model for the boring steps and a frontier model only where it earns the price.
None of that requires you to believe the Aeon rumour. It is what you would want in place on the morning it turns out to be true.

The claim underneath the claim
The most interesting thing about the Aeon story is what it is being used to explain. Through September, reports linked the name to OpenAI's internal long-horizon work on Millennium Prize mathematics — thousands of coordinating agents, an 88-hour formal-verification run, progress on a second problem. OpenAI has acknowledged progress. It has not attributed it to Aeon, and the model described in that reporting is optimised for long time-horizon reasoning, which is a different thing from a consumer agent that manages your calendar and your cart.
Two names are being welded together because they are adjacent in a config file and adjacent in a news cycle. That is the same mechanism that produced the GPT-6 Sol and Claude 5.5 entries in the signal this piece came from: a plausible model name, a real underlying event, and a version number doing more work than any evidence behind it. The correct posture is not scepticism about whether OpenAI is building an agent — it plainly is, and it has already sold the harness. It is scepticism about the name, the date, and the feature list, none of which have a primary source.
What we would tell you to do
Do not plan around Aeon this week. There is no product to evaluate, no price to compare, and no documentation to read, and the "before DevDay, maybe this week" framing has already slipped once — Altman's own walk-back moved a launch out of the week of September 14. The dated, load-bearing facts are these: "gpt-6-astra-aeon" exists as a string in a Codex Desktop feature flag as of September 3; GPT-6 Astra is OpenAI's shipping flagship at $10.00/$50.00 per million tokens with 1M context; the Agents API has been in public beta since September 10; DevDay is September 29.
Everything else is a claim. If you are evaluating agents for real work now, the products you can actually buy are Grok Bot and Muse, and the honest read on both is that neither has published independent reliability data — test on low-stakes workflows and keep production credentials out of them until someone does. If you are waiting for OpenAI's answer, the calendar gives you one date to circle and no reason to act before it.
When Aeon does ship, the first question will not be whether it is better. It will be whether it is available, at what price, and whether the monitoring that Astra ships with lets it finish a job that takes three days. Those are the numbers worth watching, and none of them are in the leak.

