
Xiaomi MiMo-V3-Flash: Is the Unreleased Fast Model Running Inside OpenCode's Omen Alpha?
- googleNEWGoogle: Gemini 3.8 Flash2026-09-0259Intelligence76Coding
- qwenNEWQwen: Qwen3.8 Max (0902)2026-09-0258Intelligence72Coding
- anthropicNEWAnthropic: Claude Fable 5.12026-09-0166Intelligence82Coding
- AlibabaNEWQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiNEWZ.ai: GLM 5.3 Flash2026-08-2658Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.15 / $0.29 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1860Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1552Intelligence68Coding
- qwenQwen: Qwen3.8 27B (free)2026-08-13qwen/qwen3.8-27b-free
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1253Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1261Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0557Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0358Intelligence72Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3152Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
- anthropicAnthropic: Claude Opus 52026-07-2463Intelligence78Coding
- googleGoogle: Gemini 3.6 Flash2026-07-2152Intelligence69Coding
- googleGoogle: Gemini 3.5 Flash-Lite2026-07-2137Intelligence49Coding
Xiaomi MiMo-V3-Flash has not been announced by Xiaomi. There is no model card, no Hugging Face weights, no pricing page and no launch post — the name is in circulation for one reason only: on September 4, 2026, teortaxesTex, a prominent independent AI analyst, said publicly that he believes Omen Alpha, the anonymous "stealth model" OpenCode began offering to its Go subscribers the same morning, is Xiaomi MiMo-V3-Flash. His identification rests chiefly on what he reads as the chain-of-thought of a DeepSeek V4 GA-era model inside Omen Alpha's visible reasoning trace — which, if it holds, would make Xiaomi MiMo-V3-Flash a Xiaomi-built fast tier trained on top of DeepSeek's open V4 weights. Nothing about that identification is confirmed, and this piece treats it as exactly what it is: one well-argued hypothesis in an active guessing game.
The ground rules for what follows: the launch of Omen Alpha is confirmed fact, and the analyst's claim is a labeled claim. The V3 generation's existence is supported by a conference preview and a leaked benchmark, both clearly marked as unconfirmed by Xiaomi. The assertion that a model named Xiaomi MiMo-V3-Flash exists at all is an inference from behavior, not a Xiaomi statement — and since Xiaomi has made none, none is repeated here as fact.
The claim, and the four reasons behind it
The trigger was a developer report. A builder testing Omen Alpha in OpenCode, posting as @SPAC89, wrote that "Omen Alpha High" — the model at its high reasoning setting — had finished a substantial coding task in under ten minutes, and that a much longer multi-hour run was under way. Quote-posting that report, teortaxesTex said he believes the model is Xiaomi MiMo-V3-Flash and gave four reasons, all behavioral rather than documentary:
• The reasoning trace. The visible chain-of-thought reads as DeepSeek V4 GA CoT — his tells are "hmm," "fine," and ✓-prefixed checklist items. The screenshot he attached shows Omen Alpha mid-build at the "high" setting with a trace that matches the pattern.
• The speed. The model is fast — that is the observation he was responding to — and Xiaomi's fast tier has historically carried the "Flash" name.
• The operator. DeepSeek does not run stealth promotions, so even if the reasoning is DeepSeek V4 GA in origin, whoever is behind Omen Alpha is a third party that built on DeepSeek rather than DeepSeek itself.
• The clustering. Xiaomi MiMo models' output text clusters with DeepSeek's — his shorthand for two generations whose text sits close in distribution, which is what a DeepSeek-trained model would look like. He names one alternative that also clusters with DeepSeek, StepFun's Step-4, and says he is unsure about StepFun's current generation.
The messenger is worth a clause, because the same fact cuts both ways. teortaxesTex is a self-described DeepSeek follower — his X profile has called himself a DeepSeek fan since 2023. That is the profile of someone who can recognize a DeepSeek V4 reasoning style on sight, and also the profile of someone inclined to see one.
Where a MiMo-V3-Flash would sit in Xiaomi's lineup
Xiaomi's officially confirmed MiMo models today are MiMo-V2.5 and MiMo-V2.5-Pro, both open-sourced under MIT at the end of April 2026. MiMo-V2.5-Pro is a 1.02-trillion-parameter mixture-of-experts model with 42 billion active parameters and a 1-million-token context window; MiMo-V2.5 is the smaller general-purpose sibling at 310 billion parameters. Behind them in the family tree sit MiMo-V2-Flash (December 2025, a 309-billion-parameter open MoE) and MiMo-V2-Pro (March 2026) — and MiMo-V2-Pro matters for this story because it first reached the public under an anonymous codename, Hunter Alpha, before Xiaomi claimed it. Anonymous field tests are not an exception for Xiaomi; they are part of how Xiaomi launches.
The V3 generation is where the record shades from shipped product into report. At ICML in July 2026, Luo Fuli, the researcher who leads the MiMo team, previewed V3's architecture. The headline is a design Xiaomi calls High Sparse: a small number of full-attention layers act as an "oracle" that selects the tokens the surrounding sparse layers should attend to, and those sparse layers share the full-attention layers' KV cache instead of keeping their own. Xiaomi says this cuts KV-cache memory by roughly ten times and validated an 11:1 sparse-to-full ratio on an 80-billion-parameter-scale model, matching full-attention quality where a hybrid sliding-window baseline lost about seven MMLU points at the same sparsity. Those are the team's own claims from a conference talk, not independent results. Then, on August 21, a purported benchmark screenshot for a MiMo-V3-Pro began circulating, shared by Max For AI: SWE-Bench Pro 72.8, against GLM 5.3's 67.9, Kimi K3's 65.8, Claude Opus 5's 74.6 and GPT-5.6 Sol Max's 75.4, plus Terminal-Bench 2.0 70.6 and τ3-bench 76.4. Xiaomi has confirmed nothing, and V3's existence and release timing remain unannounced. The takeaway for this article is narrower: the V3 generation is demonstrably in motion, but no V3 model has shipped — which is exactly why an unannounced Xiaomi MiMo-V3-Flash surfacing in stealth would be news rather than routine.
Why a Flash — and why OpenCode Go — would fit
The Flash tier is how Xiaomi seeds fast, cheap capacity; it did the same with MiMo-V2-Flash in December. And OpenCode Go is a plausible staging ground because it already runs Xiaomi's current open model: the same Go lineup that lists Omen Alpha at 11,600 requests per five hours also lists MiMo-V2.5 at 30,100, among the most generous allowances on the page. A Xiaomi team field-testing a V3-Flash next to its own MiMo-V2.5 on one platform would get a direct comparison against the previous generation at zero marketing cost, on a platform that has already run one anonymous reveal this month. That is an argument for plausibility, not evidence of identity.

The product shape points the same way. Omen Alpha at its "high" reasoning setting finishing in under ten minutes matches the behavior you would expect of a fast reasoning tier, and the third-party figures circulating for Omen Alpha — roughly $0.20 per million input tokens and $0.66 per million output, per press reproductions rather than an official rate card — put it at the cheap end of the market beside DeepSeek V4 Flash. A model that behaves like a fast, cheap DeepSeek-generation reasoner is what a "Flash" tier built on DeepSeek V4 would look like. Reported and inferred, again, not confirmed.
The DeepSeek connection is the crux — and the weakest evidence
Why DeepSeek V4 GA matters at all: DeepSeek's V4 line moved from preview to general availability in August 2026, when DeepSeek V4 Pro shipped its GA build 0813 on August 13 with MIT-licensed open weights, alongside the cheaper sibling DeepSeek V4 Flash. Open weights under MIT mean a third party may legally build on them, and the community's usual way of spotting such a model is the reasoning style it inherits: DeepSeek's GA-era chain-of-thought has recognizable verbal habits, and those habits are precisely what teortaxesTex says he sees in Omen Alpha's trace. If he is right that the trace is DeepSeek V4 GA, and right that the operator is Xiaomi, then Xiaomi MiMo-V3-Flash would be Xiaomi's fast tier built on DeepSeek's open reasoning weights — a legal but strategically notable choice, and one that would neatly explain why an ostensibly Xiaomi model reasons like a DeepSeek.
Now the hard part. A chain-of-thought resemblance is the weakest class of model fingerprint, because it is style rather than structure. The anonymous-model identification that actually held up this month — Ox Alpha as GLM-5.3-Flash — was settled on much harder evidence: tokenizer matching and video-encoding forensics. No comparable hard fingerprint has been published pointing Omen Alpha at Xiaomi. The community's leading read, moreover, points elsewhere: developers report that Omen Alpha's backend requests resolve under a "zhipu" namespace, some testers describe the model as self-identifying as GLM-based, and the press reading of the launch is that Omen Alpha is Zhipu AI's second anonymous release in nine days, following the Ox Alpha run that was read as GLM-5.3-Flash. If that reading is right, the DeepSeek-flavored reasoning is either a stylistic coincidence or a sign that the Zhipu model was itself trained with DeepSeek traces — but it is not a Xiaomi product. The two camps cannot both be correct, and neither camp has a statement from the lab it names.

What would settle it
• A Xiaomi statement — or weights. Xiaomi has open-sourced its MiMo models since MiMo-V2-Flash, with MiMo-V2.5 and MiMo-V2.5-Pro released under MIT; a real Xiaomi MiMo-V3-Flash would most plausibly be confirmed the way GLM-5.3-Flash was, with weights and a model card on Hugging Face.
• A hard fingerprint. Tokenizer matching was the evidence that identified Ox Alpha; a tokenizer test pointing Omen Alpha at Xiaomi has not been published.
• The model's own account. Some testers report Omen Alpha describing itself as GLM-based. A model's self-report is data, but it is not provenance.
• Independent benchmarks of Omen Alpha. Neutral evals of the anonymous model would let the community compare it against the MiMo-V3-Pro numbers that leaked in August.
• The reveal after the Go window. Past anonymous models were claimed days to weeks after launch; whoever claims Omen Alpha will name the model exactly.
What to do while the guessing continues
The practical position does not depend on who is right. Omen Alpha is an anonymous model behind a $10-a-month subscription, and Xiaomi MiMo-V3-Flash is an unconfirmed name attached to it; neither is something to bet a production pipeline on today. When a maker does step forward and the underlying model reaches ordinary APIs, that is the moment a routing layer earns its keep: OrcaRouter passes provider list prices through at 0% markup, so a newly published rate is live on the same API the day the vendor publishes it, and automatic failover lets you send a slice of real traffic at an unproven newcomer with a fallback to a model you already trust. Meanwhile the thing to watch is simple — does anyone claim Omen Alpha, and does Xiaomi MiMo-V3-Flash ever appear on a model card? One of those two will happen.

As of September 4, 2026, Xiaomi MiMo-V3-Flash is a name, a plausible product shape, and a prominent analyst's hypothesis — not a released model. The V3 generation is real enough to preview in July and leak in August. Whether its first public sighting is happening right now under the codename Omen Alpha is a question only Xiaomi can answer, and it has not answered it.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
