
Space Bunny: The Reference Page for the Stealth Model Also Listed as Space Bunny Alpha
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 517 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 196 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Space Bunny and Space Bunny Alpha are the same model. Settle that first, because the two names pull in different directions: the bare name is what most people type, and the versioned name is what every listing and every write-up uses. There is no second Space Bunny — no smaller sibling, no earlier release the Alpha supersedes. The bare name is the short form; Alpha is the preview-stage marker this family of anonymous listings carries while the operator stays unnamed. Space Bunny is a stealth preview served on a third-party platform: one million tokens of context, a 524,288-token output ceiling, text, image and video input, and a listed price of zero dollars per million tokens, with the provider's identity undisclosed. This page is the reference for it — the destination for someone who searched the model's own name, not a second announcement that it exists.
Two things about how to read what follows. Everything under the envelope heading is operator-supplied listing metadata: the anonymous provider's own numbers, published on the platform hosting the preview and re-read today, 28 September 2026. None of it is independently verified, and I say so again where the numbers appear. The benchmark section is separate and sourced third-party throughout. The other note is why this page exists at all. The listing dates from 23 September 2026, five days ago, so there is no age problem to argue around — but the warrant here is not that date. It is first-party search demand we measured ourselves: 1,668 impressions and 238 clicks on the query "space bunny" in the 28 days to 25 September, at position 4.8. That is a fact about what people type, not about the model. Readers were already clicking through to whatever page we happened to rank; there was simply no page about the model itself for them to land on. Filling that gap is the clause this piece stands on — standing demand measured on 28 September 2026, with the listing date used to date the model rather than as an event to report.
The name, settled properly
The evidence for "same model" is an absence, so it is worth being precise about what was checked. The platform hosting the preview lists exactly one entry for the family, under the identifier stealth/space-bunny-alpha/ and the display name "Space Bunny Alpha"; there is no bare space-bunny/ entry beside it. The model's own field guide — a third-party site that states plainly it is not affiliated with the platform or the developer — introduces it as "Space Bunny Alpha, also written space-bunny-alpha" and dates its listing to 23 September 2026. The Hugging Face organisation that hosts companion repositories for this series carries one bunny repository, named for the alpha. Looking for a bare "Space Bunny" as a separate model turns up nothing but the same model under its shorter name.
Practically: treat the two strings as aliases. "Space Bunny" is the head term and the one with the better click behaviour; space-bunny-alpha/ is the identifier you put in a request. The one thing the bare name does not imply is a version — there is no documented 1.0, no smaller variant, no non-preview release. A listing or a post that implies otherwise is inventing a family the record does not show.
The suffix does carry real information: the model is in preview, operated by a provider who has chosen not to be named during it, and the platform's stealth notice says prompts and completions may be retained by that provider but are not used for training. Read that clause carefully, because it is a retention statement and not a zero-retention one. If you have read our pre-launch coverage of this listing — the Space Bunny Alpha leak page — that page's function is served: what it anticipated is now simply the model's state, and the live questions have moved from identity to envelope, cost and throughput.
The envelope, as the listing states it
Re-read today on the live listing. Operator-supplied, unreproduced.
• Context window — 1,000,000 tokens. The same 1M class as the previous anonymous previews in this series.
• Maximum output — 524,288 tokens. Half the context window, which is a very large ceiling at any price and one no third party has stress-tested for coherence at the far end.
• Input and output — text, image and video in; text out. Video input is the detail that most distinguishes this generation of anonymous listings from ordinary multimodal releases.
• Reasoning — mandatory and cannot be switched off, with five effort levels: low, medium, high, xhigh and max. The live listing today shows the provider default as max, and the third-party field guide independently records the same default. Hiding the reasoning text through the API does not disable the computation; it still consumes output budget.
• Tool use and structured output — tools/ and tool_choice/ are accepted, so function calling is available. response_format/ gives JSON output without JSON-schema enforcement, the same caveat this family of previews has carried throughout. If your pipeline needs a validated object back, validate it yourself.
• Not disclosed — provider name, parameter count, active parameters, architecture, tokenizer family (filed only as "Other"), knowledge cutoff, quantisation, licence, and whether releasable weights exist at all. The companion repository says GGUF files will appear "once the model weights become available" — a promise about a future file, not evidence that weights exist now.
• Operational, and notably thin — the listing publishes no throughput figure and no latency figure. Both fields are empty. It reports a single serving provider, tagged as stealth, with unknown quantisation.
That last bullet is the honest answer to "how fast is it": nobody who published a number owns the endpoint, and the endpoint's operator has published nothing. That is not a reason to avoid the model. It is a reason to measure it on your first call rather than trust an adjective in a description.

The rate card: zero, and what zero means here
The listing prices the model at $0 per million tokens in and $0 per million out — a free preview, on the operator's own listing, read today. No per-request limit and no rate-limit policy are published on the entry.
Two qualifications. First, that is the operator's listed price today, not a commitment: a stealth preview price is a property of a listing, and listings change without notice. Second, in this series a $0 preview has historically been an evaluation instrument rather than a permanent tier — the free window has closed when a preview ended. Anyone building on a zero-dollar rate should treat the number as current rather than durable, and should know what their per-million fallback costs before they need it.
The endpoint surface
What a caller actually gets, taken from the listing's supported-parameter set rather than from prose documentation.
• Chat completions — messages/ with optional stream/, so ordinary streaming generation.
• Sampling and limits — max_tokens/, temperature/, top_p/. Reasoning tokens come out of the max_tokens/ budget, so a tight cap truncates thinking before it truncates output.
• Reasoning controls — reasoning/ and reasoning_effort/ across the five-step ladder, plus include_reasoning/ to surface or hide the reasoning text.
• Structure and tools — tools/, tool_choice/ (auto is supported; forcing a specific function is not), and response_format/ for JSON without schema validation.
That is the whole surface. No separate vision endpoint, no batching API, no fine-tuning path, no embeddings. Multimodal input goes in through the same chat call, which is convenient and also means a large video attachment competes for the same context budget as everything else in the conversation.
What third parties have actually measured
This is where most search results stop being careful, so provenance first. Artificial Analysis, the independent leaderboard we normally cite for cross-model comparison, has no entry for Space Bunny under either name — checked today by direct request. The two third-party measurements that do exist are different in kind, and neither is a matched harness against the models it is plotted beside.
The first is AI BENCHY, a third-party benchmark site publishing per-model runs. Its entry for Space Bunny Alpha records 7.0 out of 10 at high reasoning, a 62.1% attempt pass rate, 10.0 reliability, an average response time of 27.38 seconds and $0.000 recorded cost, and places it at #161 in that site's own table. By category, its strongest results are data extraction and tool calling, both 10.0 out of 10; its weakest are trivia at 3.0 and general intelligence at 4.2. Those are that site's figures from its own suite. The number worth reading is the pass rate and the category spread, not the global rank, which is a position in somebody else's ordering.
The second is the model's own field guide, which publishes its own evaluations: 82.0% on a 60-question GPQA Diamond subset, 75% on MMLU-Pro, and 46.1% on a 300-question subset of Humanity's Last Exam with an approximate 95% confidence interval of 40.4–51.8%. The site is explicit that these are its own subset runs and that the models shown beside them are published reference results from their respective vendors — different question counts, different harnesses, different setups. That caveat is the important part and should not be quietly dropped: those columns are not a head-to-head, and nobody has run a matched-effort, matched-harness comparison of Space Bunny against any model it is plotted next to. A page that presents those numbers as a ranking is doing something the source itself declined to do.
Two further third-party results are worth having because they test things benchmarks miss. The field guide reports a long-context check in which three distinct codes hidden inside a single input of roughly 200,000 tokens were all returned in the correct order on one trial — a capacity test rather than a capability score, and the kind of check a 1M context window invites and rarely gets. It also reports a token-efficiency comparison in which Space Bunny consumed 305,989 output tokens against 913,989 for Qwen3.8 Flash on the same benchmarks: about two thirds fewer, which matters more than it sounds once the free window closes.
The clearest single fact in the whole picture is a gap. There is no Artificial Analysis entry, no arena rating, and no benchmark reproduced by a second party. One 7.0/10 from one suite and one 46.1% from a self-published subset, with no replication anywhere, is a starting position rather than a verdict.

Who is behind it
Undisclosed, and the listing says so itself: the model is developed and operated by a third-party provider who has chosen to remain anonymous during the preview, and the platform hosting the listing states that it is not the developer, owner or provider. The companion repository carries no lab name, no licence and no weights.
There is an active public guessing discourse — prompt-format reconstructions, tokenizer-family probes, sample comparisons — and none of it has produced a confirmed identity. We are not repeating the hypotheses, for the same reason we did not when this family's previous entries were anonymous: a shared benchmark score or a matching token count is a resemblance, not an identification, and a community guess dressed as a finding is worse for a reader than an honest blank. What the record does show is that every anonymous preview in this series so far has eventually been claimed by a named lab, usually within days to weeks and often with a price change attached. That is a base rate about how these previews end, not a prediction about this one — and it is the reason to treat today's envelope as a snapshot rather than a specification.
Evaluating an unknown model without betting a path on it
An undisclosed operator does not stop you testing the model. It changes which questions you can answer. Three things follow from the listing itself.
• The retention clause is the one to respect. Prompts and completions may be retained by a provider with no name. Public benchmarks, open-source code and disposable prototypes are fine; customer data, proprietary source and anything under a compliance regime are not, and a free price does not change that arithmetic.
• Treat every capability claim as a hypothesis until you have your own numbers. Throughput, first-token latency and long-output coherence are the three figures nobody has published for this model, and they are also the three that decide whether it fits a workload. They cost an afternoon to measure.
• Do not make it the only model in a path. A provider with no name, no status page and no escalation route is a provider whose failures you cannot chase.
That last rule is where a router earns its place, and it is the part we can speak to directly. OrcaRouter serves 200+ models behind a single API with no markup on provider list price — you pay the provider's own published rate, so a vendor price cut is live on our side the same day — plus automatic failover across providers and a routing DSL that lets you declare a known-good model as the fallback behind an experimental one. To be exact rather than generous about the availability question: Space Bunny is not on our catalogue today. Our model API returns no such route under either name, so we are not a way to call this model, and this page is not a route to it. What we are is a way to keep a measured model behind an unmeasured one while you decide — which is the only sensible posture toward a free preview from an operator nobody can name.
What would change this page
Four things, in the order they are likely to arrive. A named owner, which in this series has come with a licence, a parameter count and a price. A second-party benchmark reproduction, which would be the first quality evidence not tied to a single suite. Published throughput and latency, the two empty fields a buyer would most want filled. And a change to the rate card, currently the model's single most decisive property.
Until then the honest summary is short. Space Bunny is a free, one-million-token, video-capable reasoning preview with an undisclosed operator: one 7.0/10 third-party run, one set of self-published subset evaluations, no independent leaderboard entry, and no throughput figure at all. Its bare name and its alpha name are the same model. Everything else about it is a date-stamped snapshot of a listing, and this page is the place to check which snapshot is current.

