A hero title card for Ox Alpha, the anonymous frontier model, showing a question-mark-shaped model chip with text, image and video inputs flowing in and token blocks flowing out, with three pills reading '1M token context', 'Free for one week', and 'Maker unknown', and the OrcaRouter logo composited in the bottom-right corner.
Guides & Insights

Ox Alpha: The Anonymous Frontier Model That's Free for a Week — and Nobody Will Say Who Made It

Author

Magnus Corvin

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On August 20, an anonymous frontier model called Ox Alpha appeared on a third-party AI API platform with a 1-million-token context window, text/image/video input, and a price of zero. Nobody has said who built it — and that silence is itself the story. The last four models released the same anonymous way were all eventually claimed by Chinese labs: Zhipu AI's GLM-5, Xiaomi's MiMo-V2-Pro, Ant Group's Lingxi Ling-2.6-flash, and Meituan's LongCat-2.0. Ox Alpha, listed under the ID stealth/ox-alpha, is the fifth entry in that series. It is free for the next week, and on the pattern of the previous four, the reveal usually comes fast.

Every figure in this piece is the operator's own claim, data from the platform's own listing, or a public announcement — there are no independent benchmarks yet, and nothing here should be mistaken for one.

What we know — and what we don't

The fast version:

• The operator describes Ox Alpha as "a frontier model built for efficient coding, sustained agentic work, and real-world production use" — a reasoning model aimed at long-horizon software engineering and workflows that mix text with visual context.

• It has a 1,048,576-token (1M) context window, a 131,072-token output limit, and takes text, image, and video input — the first of the anonymous releases to advertise video input.

• It is free for one week: $0 per million tokens in and out, with the operator claiming "generous rate limits, near-unlimited usage" and 100 trillion tokens per day of serving capacity.

• Early activity data on the platform already shows real production traffic, including a coding agent from Nous Research and the Zed editor.

• No company has claimed it as of August 21, there are no benchmark scores of any kind, the real price after the free week is unknown, and the model's architecture and parameter count are undisclosed.

What Ox Alpha is

The platform's listing describes Ox Alpha as "a reasoning model designed for coding, sustained agentic work, and production workloads — long-horizon software engineering, complex reasoning, and workflows that combine text with visual context." That is a specific positioning: it is built for long-running agent loops and for code, not for general chat. The 1M context is the give-away — that is enough room for a multi-hour agent session or a large repository — and the 131K output cap allows for long single generations. It supports tool and function calling and structured JSON output, which is the rest of the agentic checklist.

A single-column scoreboard for Ox Alpha listing: context window 1M (1,048,576) tokens, max output 131,072 tokens, input text/image/video, independent benchmarks none yet, price free for one week, throughput 56 tokens/s (P50, platform-measured), with a footer reading 'Operator + platform-reported; no independent scores yet', and the OrcaRouter logo in the bottom-right corner.

The video input is the most distinctive item on the sheet. None of the earlier anonymous models advertised it, and a model that can take video frames as well as text and images is aimed at a different class of workload than the chat-tuned releases that preceded it. All of this — the description, the specs, the speed figures — comes from the operator and the platform's listing. None of it is independently verified.

The free-for-a-week deal

The promotion is aggressive. Free for the next week, "generous rate limits, near-unlimited usage," and a claimed 100 trillion tokens per day of capacity, with the operator's note inviting users to "see what you can do." Taken literally, 100T tokens a day would put this operator among the largest inference providers anywhere — the claim reads less like a spec and more like a stress-test challenge, which fits the pattern: anonymous releases are how a lab gets frontier-scale real-world traffic without putting its name on the door.

The other side of "free for a week" is that there is no known price after the week. The earlier anonymous models either stayed free or were folded into their reveal later; Ox Alpha's free window creates a deadline, and the real number could land anywhere. That is the main reason not to design a production budget around the current $0.

No benchmarks yet — but already real traffic

Ox Alpha's listing has no benchmark data: no intelligence index, no coding score, no agentic score, and it does not appear on independent leaderboards such as Artificial Analysis. The word "frontier" is the operator's own, and there is no verified evidence yet of how good the model actually is.

A screenshot of the Artificial Analysis models leaderboard as of August 21, 2026, showing the Intelligence section led by Claude Opus 5 and Claude Fable 5, the largest context-window highlights, and the output-speed rankings — a frontier leaderboard Ox Alpha has no independent score on yet.

What does exist is usage. The platform's activity data shows real production traffic within the first hours — including Hermes Agent, the agentic coding project from Nous Research, and the Zed editor. Both are exactly the "sustained agentic work and coding" use cases the model is positioned for. It is not a score, but it is a signal: developers with real workloads decided a free, frontier-class, anonymous gamble was worth wiring up. For a model with zero published benchmarks, that early adoptership is the only evidence there is.

The data-retention fine print

The free-week announcement touts "zero data retention." The model's listing on the platform tells a more qualified story: prompts and completions are retained by the provider but not used for training, under the platform's stealth-model terms. The difference matters if you are about to paste proprietary code into a 1M-token window owned by an anonymous operator. Treat "zero data retention" as marketing until the operator publishes terms that say so in so many words — and route anything you would not want logged through a separate, attributable model.

The anonymous-model playbook

Ox Alpha is the fifth anonymous release in six months, and the previous four resolved the same way — an anonymous debut, a burst of free traffic, then a company stepping forward:

• Pony Alpha (February 2026) — confirmed by Zhipu AI about five days after launch as GLM-5, its 744B-parameter MoE flagship.

• Hunter Alpha (March 11) — a free, trillion-parameter model with a 1M context that became the speculation story of the spring, widely guessed to be DeepSeek V4; revealed as Xiaomi's MiMo-V2-Pro, with the companion Healer Alpha identified as MiMo-V2-Omni.

• Elephant Alpha (April) — a free, efficiency-focused text model that Ant Group claimed as Lingxi Ling-2.6-flash about two weeks after it appeared.

• Owl Alpha (late April) — an agent-focused model with native tool calling and a ~1M context that Meituan confirmed as LongCat-2.0 on June 30, the first trillion-parameter model trained and served entirely on domestic Chinese chips.

• Ox Alpha (August 20) — unclaimed as of this writing.

A two-column scorecard titled 'The anonymous models — and their reveals', listing Pony Alpha revealed as GLM-5 by Zhipu AI, Hunter Alpha as MiMo-V2-Pro by Xiaomi, Elephant Alpha as Lingxi Ling-2.6-flash by Ant Group, Owl Alpha as LongCat-2.0 by Meituan, and Ox Alpha marked '??? — unclaimed', with a footer noting reveals per each lab's public announcement and Ox Alpha unclaimed as of August 21, 2026, and the OrcaRouter logo in the bottom-right corner.

Why launch a good model behind a mask? The standard reading, which the Hunter Alpha cycle made concrete, is that anonymity removes brand bias from evaluations, and the free usage crowdsources real-world evaluation data at frontier scale while everyone argues about the owner. As one analysis of the pattern put it, "the lack of a brand is the marketing." The extra win is speed: an anonymous model can reach a global developer audience in hours without a launch event, and the mystery itself becomes the distribution.

Who is Ox Alpha?

No one has claimed it, and there is no hard evidence yet. What exists is a prior: four of four resolved anonymous releases turned out to be Chinese labs, and Ox Alpha's profile fits the same shape — a coding- and agent-focused reasoning model with a 1M context and a confident capacity claim is exactly the kind of thing a major vendor ships to test its serving stack at scale.

The distinguishing signal is the video input, which none of the earlier anonymous models advertised; that narrows the field toward labs with serious multimodal and video work. But a prior is not evidence, and this release differs from its predecessors in one way worth noting: the operator is claiming a no-retention posture, whereas earlier anonymous models were explicit about collecting your data. If that holds, the motive shifts from data-flywheel to pure scale-testing or pre-launch positioning. Until a company steps forward — or the model leaks an architecture tell — every guess is a guess.

Should you build on it this week?

This week it costs nothing but integration time. If you do long-horizon agentic work, code on long contexts, or run video-heavy pipelines, that is exactly the intersection where Ox Alpha is positioned — and the free window is the right time to find out whether the positioning is real. Two cautions. First, there is no independent evidence it is frontier-class; the "frontier" label is a claim. Second, there is no guarantee the free tier — or the model — survives the week. That is the shape of a test, not a dependency.

The production-safe way to run a test like that is a fallback chain: the experimental model answers when it is healthy, and the request rolls to a proven model the moment it stalls, errors, or disappears. That is what a routing layer is for. One API across 200+ models, automatic failover so an anonymous model's free-week launch traffic does not become your outage, and 0% markup — provider list price passed through — so that when Ox Alpha's real price appears after the free week, the number you see is the number you pay, the same day the operator changes it. Use the free week to audition; decide after the reveal.

What to watch

Four things will tell the story. The reveal: past anonymous models were claimed days to weeks after launch, and when a company steps forward the model page will name it. The price: whatever replaces the free tier, and whether a free tier survives at all. The first independent benchmarks: an Artificial Analysis or LMArena entry would settle whether "frontier" is a fact or a tagline. And the capacity: whether the claimed 100T tokens/day holds up when real load hits it.

The honest verdict: Ox Alpha is an unusually cheap experiment in both directions. The platform tests whether an anonymous frontier-class model at $0 can win real production traffic in a week; you get to test whether the model deserves any. Run the experiment, but run it with a fallback underneath. By the time the free week ends, the odds are good that we will know who made it, what it costs, and whether it earned a place in your stack — and those three answers are what make the audition worth having.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube