
Muse Code: Who Makes It, What It Costs, and Which Version Is Current
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 517 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 196 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Muse Code is Meta's coding agent for the terminal and CI, built for the Muse Spark model family, and the two versions that matter to anyone arriving here are Muse Spark 1.3 — the current generation, dated September 2, 2026 — and Muse Spark 1.2, which is the version the CLI still selects unless you tell it otherwise. Muse Code installs as a native muse binary, runs interactively or headlessly, and is sold three ways: three monthly plans at $5.00, $15.00 and $50.00, and pay-as-you-go token billing on Meta Model API at $1.25 per million input tokens and $4.25 per million output tokens on the Standard tier. One generation back again, Muse Spark 1.1, is still served on the same tier at the same rates.
Read this as a reference page, not as news. Muse Code itself is not a new thing to report: Meta shipped it as a beta on August 5, 2026 and took it out of beta on August 31, 2026, and this page uses those dates to place the model in time rather than to announce it. The warrant for the page is standing search demand that we measured first-hand on September 28, 2026, where the bare query "muse code" drew 4,547 impressions in the trailing 28 days with zero clicks, and the commercial spelling of the same intent — "muse code free", "muse code plan", "muse code plans", "muse code free tier", "muse code 1.3" — drew another 124 impressions, also with zero clicks. Readers are asking four specific questions. This page answers them in order, from Meta's own documentation, and says plainly when a figure comes from Meta rather than from an independent evaluator.
If what you actually need is a review of the agent's behaviour, or a head-to-head against another coding harness, this is not that page. Those already exist on this blog and the ones worth your time are linked at the end.
What Muse Code is, and who makes it
The vendor definition is one sentence: "Muse Code is Meta's coding agent for the terminal and CI, built for Muse Spark." It is not a library, not a chat interface, and not a model — it is a harness. You run it inside a project directory; it plans, edits files and runs commands to complete a task, with approvals and an OS sandbox active from the first run.
Meta frames Muse Code and Meta Model API as two routes to the same underlying model: call the API directly when you are building your own agent or application, run Muse Code when you want a ready-made coding agent at the command line or inside a pipeline. That distinction matters for pricing, because the subscription buys the harness and the API key is billed separately per token.
Installation is a one-line installer, and the two vendor surfaces do not quite agree on the string:
• macOS and Linux — curl -fsSL https://dev.meta.ai/install.sh | sh on the docs page, and the same command ending | bash on the marketing page and in both research blog posts
• Windows — irm https://dev.meta.ai/install.ps1 | iex, published on the docs page only
• Verification — muse --version, then muse in a project directory to start an interactive session
• First run — you are asked to trust the workspace (which is what loads its skills, rules and hooks) and to authenticate either through a browser sign-in or with an API key; in CI you set META_API_KEY instead
• Two surfaces — muse for the interactive terminal UI, muse exec "<prompt>" for a single non-interactive run to completion
• Drive it from your own program — muse serve runs a versioned session protocol, muse schema prints its JSON schema, and npm install @muse-code/sdk installs the TypeScript wrapper
That last point is the part most summaries miss. Muse Code is not only a CLI: the session protocol is a documented, versioned interface, and the changelog shows it being extended deliberately — sessions can be renamed over the protocol, clients can read a structured diff summary for each edit, and a driving program can set a session-wide reasoning-effort default.

What Muse Code costs, and the free-tier question answered plainly
There is no free tier. That is worth stating flatly because it is the single most-asked commercial query in the demand data and the answer is unambiguous: the Muse Code product page, the subscriptions page and the changelog together contain no free plan, no trial, no free credit allocation and no rate-limited gratis access.
One caution on that check. Meta's pricing documentation does use the words "platform free-tier credits" once — in the Muse Voice Transcribe note, which says zero-data-retention transcription is priced at parity with the Standard tier and that platform free-tier credits apply. Muse Voice Transcribe is a separate product on Meta Model API with its own per-hour billing; that sentence says nothing about Muse Code. If you find that line quoted as evidence of a Muse Code free tier, it is being read out of context.
What Meta does publish is three subscription tiers, all billed monthly, plus pay-as-you-go token billing for anything beyond the plan's prompt allowance:
• Everyday Usage — $5.00/month; access to the latest Muse models; 10–50 prompts every 5 hours, including image and video uploads; voice mode; web search
• High Usage — $15.00/month; everything in the Everyday plan; 5× more usage than Everyday; more prompts; more multimodal inputs
• Power Usage — $50.00/month; everything in the High Usage plan; 20× the Everyday usage allowance; expanded prompts; early access to new features; higher file uploads
• Standard token rates — $1.25 per million input tokens, $4.25 per million output, $0.15 per million cached input, for muse-spark-1.3, muse-spark-1.2 and muse-spark-1.1 alike
• Contributor token rates — $0.10 / $0.20 / $0.002 per million input, output and cached input, for muse-spark-1.3-contributor and muse-spark-1.2-contributor
The Standard and Contributor tiers differ in exactly one substantive way, and Meta states it on both the models page and the rate card: on Standard, your prompts and completions are not used to train Meta models; on Contributor, they are. The product page labels the Contributor rows "Used to improve our products" and the Standard rows "Not used to improve our products." Contributor is roughly twelve times cheaper on input and about twenty-one times cheaper on output, which is a large enough gap that it is a data-governance decision before it is a budget decision.
Three pricing details are easy to miss and are all worth knowing before you size a workload:
• No long-context premium — Meta's rate card says this explicitly: you pay the same rate whether the context window is mostly empty or almost full, which is unusual and materially changes the economics of long agent runs
• Web search grounding is metered separately — $2.50 per 1,000 search queries on top of the request's token cost, and it applies to text models such as Muse Spark
• Rate limits are per team, not per key — 3,000 requests and 4,000,000 tokens per minute on Standard, 100 requests and 3,000,000 tokens per minute on Contributor; using several keys inside one team does not multiply the quota
Subscription mechanics are documented on Meta's subscriptions page and are conventional: an upgrade is prorated and immediate, a downgrade takes effect at the next billing cycle, cancellation requires at least 24 hours before the billing date, and there are no refunds for a cancelled subscription except where the law requires one. Subscriptions cover the Muse Code CLI credential specifically; additional API keys bill pay-as-you-go.

Which version is current: 1.2.1, Muse Spark 1.3, and the default that is neither
This is the question the demand data is loudest about — the query "muse code 1.3" exists at all — and it is genuinely confusing, because "the version" can mean three different things and Muse Code publishes all three.
• The CLI release — 1.2.1 is the current one. Meta's changelog lists its releases as 1.2.1, 1.1.1, 0.2.1 and 0.1.0, the last of which is labelled "Launch version"
• The model generation — Muse Spark 1.3, announced September 2, 2026 and described on Meta's models page as "the latest version … Recommended for new work"
• The CLI's default model — muse-spark-1.2, stated in the first-run section of the docs and again under model selection in the configuration page
So a fresh install of Muse Code 1.2.1 will run Muse Spark 1.2 unless you pass --model muse-spark-1.3 or switch mid-session with /models. Meta's own quickstart pages use muse-spark-1.3 in every API example, which makes the split easier to misread than it should be: the API examples and the CLI default do not point at the same model today.
The 1.2.1 changelog describes what changed in the harness rather than in the model, and the headline items are:
• Voice input on by default on macOS, bound to Option+V and managed with /voice
• A /rewind command that returns the conversation to an earlier input, sharing the picker used by double-Esc
• [Image #N] labels on pasted and dropped images so the model can be asked about a specific one by number, with labels and source paths surviving resume, rewind and forks
• A bundled migrate skill that imports memory notes and MCP server definitions from Claude Code or Codex into Muse Code
• /mcp for a live inventory of connected MCP servers and their tools
• New sessions opening in the Auto-review permission profile, which grants the same access as "Ask me" but routes eligible approval requests to an automated reviewer, falling back to asking you when the reviewer is unavailable
• Security fixes, three of which are behavioural rather than cosmetic: commands launched through wrappers such as env and setsid are now reviewed as the command they actually launch; a mid-session permission-mode change applies to already-running tools at their next action; and the Unrestricted permission profile now behaves exactly like --yolo
On the model side, the September 2 announcement is explicit that Muse Spark 1.3 was trained across a diverse set of harnesses, and that it was the reasoning tier — "Muse Spark 1.3 with max reasoning is now available on Muse Code and Meta Model API" — rather than the base model alone that Meta was promoting. The Muse Spark 1.2 launch post from August 5 adds the other half of that story: Muse Spark 1.2 was co-trained with Muse Code itself, using rejection-sampled harness trajectories plus recipe work on goals, compaction and subagents, with the Muse Code toolset integrated to maximise harness compatibility. The model and the harness were developed against each other.
The envelope: context, modalities, platforms
Muse Spark 1.3, Muse Spark 1.2 and Muse Spark 1.1 all share one envelope, and Meta states it in a single table row: 1,048,576 tokens of context, with text, image, video, audio and PDF as inputs and text as the only output, across every tier and every version.
Modality parity comes with one caveat that Meta publishes as a footnote and that is worth repeating rather than burying: audio understanding in Muse Spark 1.3 is not fully supported and response quality for requests that include audio content may be degraded; Meta's own guidance is to use Muse Spark 1.2 for audio, or Muse Voice Transcribe for dedicated speech-to-text. In the CLI that is actionable — --model muse-spark-1.2 is a documented override, and the default already points there.
Reasoning effort is a harness-level setting with eight levels, and Meta's configuration page names them in order: none, minimal, low, medium, high (the default), xhigh, max and ultra. Three details in that list are load-bearing:
• max is the deepest reasoning level, described as extended reasoning beyond xhigh, and Muse Code offers it on both the Standard and Contributor tiers
• ultra is a client-side setting rather than a model tier: it maps to each provider's highest supported reasoning tier and can make Muse Code delegate more aggressively, and where a provider does not support it a request for ultra runs at xhigh
• The Meta provider does not accept none, so the lowest level is available only through other providers configured in the harness
Platform support is where Meta's two surfaces disagree, and the disagreement is small but real. The docs page says Muse Code runs on macOS, Linux and Windows from one codebase, and lists three Windows-specific behaviours: PowerShell replaces Bash, so agent-run commands use PowerShell syntax and Windows command-line tools; the sandbox may prompt for administrator approval, and Windows can show a User Account Control prompt the first time the sandbox initialises, though normal use does not require running as administrator; and two capabilities are unavailable on Windows — voice input and session messaging, both of which work on macOS and Linux. The product page, by contrast, says only "Available for MacOS and Windows." Both statements come from Meta's own site today; read together they mean Windows is supported with documented gaps, and Linux is supported even though the marketing page does not mention it.
One further availability nuance comes from the docs rather than the marketing copy: Workflows — the parallel and staged agent orchestration — coordinate work only when the installed build includes the workflow engine and the rollout is enabled, and Meta says plainly that they "are not available on every build or platform." If a workflow feature is the reason you are installing, check that your build has it rather than assuming it.
What Meta claims, with the harness and effort named
Meta publishes benchmark material, so the honest framing here is not "no vendor benchmarks" — it is "these are vendor benchmarks, with the comparison conditions stated where Meta states them."
The single quantitative claim Meta makes in prose is in the Muse Spark 1.3 announcement, and it is attributed to internal comparison rather than to a third-party evaluation: "In comparisons by Meta engineers, it proved to be significantly faster and more efficient, using ~20% fewer tool calls and ~25% fewer tokens." The baseline is stated — relative to Muse Spark 1.2 — and the effort level is in the announcement's own framing, since the release is titled around Muse Spark 1.3 with max reasoning. What is not stated is the task set, the number of runs, or the harness configuration beyond the fact that it is Muse Code. Treat the two percentages as a vendor-reported direction of travel, not as a measured benchmark result.
The other published result is a case study rather than a score, and it is the more interesting of the two. Meta reports testing Muse Spark 1.2's ability to iteratively optimise GPU kernels over 1,000+ tool calls and up to 24 hours, using Muse Code's agentic coding environment to write, compile, profile and progressively improve kernel performance against a provided baseline. The workload is named — KDA and MLA kernels on NVIDIA Hopper GPUs — and so is the constraint: third-party kernel libraries such as FLA were prohibited, and the baseline was a Triton FLA KDA implementation. A 24-hour, thousand-tool-call run is a different kind of evidence from a benchmark percentage. It says something about endurance under a long-horizon objective and very little about comparative quality.
The launch charts for both Muse Spark 1.2 and Muse Spark 1.3 plot those models against a small set of named competitors across agent, coding, instruction-following and long-context evaluations — DeepSWE v1.1, SWE-Atlas-QnA, Terminal-Bench 2.1 and 4.0, tau2-bench, GDPval, SciCode, IFBench, MultiChallenge and AA-LCR v1.1 appear among the axes. Those charts are vendor-produced and vendor-plotted, and the competitor configurations on them are Meta's choice, not a neutral panel. We are not reproducing the bar values here for that reason: a vendor chart read as a leaderboard is exactly the failure mode this page is meant to avoid.
Two honest gaps are worth stating rather than glossing:
• Meta does not publish a head-to-head of Muse Code against Muse Spark 1.2 at matched effort on a named task set; the ~20%/~25% figure is the closest thing, and it is an internal comparison with stated efficiency deltas rather than scores
• Meta does not publish a live-latency evaluation of the CLI, and none of the published numbers describe what a session feels like under real editing load
What Artificial Analysis measures independently
Artificial Analysis is the independent evaluator that currently publishes a coding-agent measurement covering Muse Code, and it is a different kind of number from anything above: the rows are harness-and-model pairs, evaluated by AA rather than by the vendor.
On the Artificial Analysis Coding Agent Index — the page identifies itself in its own structured data as v1.5 — the Muse Code rows read:
• Muse Code running Muse Spark 1.3 at max reasoning — 54.3
• Muse Code running Muse Spark 1.3 at xhigh reasoning — 48.3
• The top of the same board — Claude Code running Claude Opus 5.5 at max reasoning, at 66.0
Two things about those figures matter more than the figures themselves. First, the 5.0-point gap between the max and xhigh rows is the same model in the same harness with one setting changed, which is a useful reminder that a "Muse Code score" is underspecified unless the effort level is named. Second, this board is a different measurement from the one our own September coverage of Muse Spark 1.3 quoted: the earlier article cited 68 for Muse Spark 1.3 at max, 64 at xhigh, and 68 for Claude Opus 5 at xhigh. Those numbers no longer appear on the live board — the current top row is Claude Opus 5.5 at max with 66.0, and the Muse Code rows read 54.3 and 48.3. The index has been revised and the composition of the field has changed. If you are carrying the older figures around, they are superseded, and the two sets should never be quoted side by side as though they were the same measurement taken twice.
The AA model page for Muse Spark 1.3 is a separate surface again, and its own figures are AA-measured rather than vendor-reported: an Intelligence Index of 48.09, Terminal-Bench 2.1 at 84.3%, GPQA Diamond at 93.5%, HLE at 48.7%, SciCode at 58.8% and long-context recall at 83%, with a release date of 2026-09-02, a 1,000,000-token context window and classification as proprietary, closed-weight. Note that this is one index revision on one model page — AA publishes per-model indices that are revised on their own schedule, so mixing a figure from this page with a figure from a different model's page does not give you a same-snapshot comparison.
What none of this gives you is a matched comparison. The AA coding rows cross harnesses by construction — Muse Code against Codex, against Claude Code, against Grok Build, against Kimi Code CLI — so a row-versus-row reading compares a harness-model pair to another harness-model pair, not two models under one harness. Where Meta has not run the matched comparison and AA's board is not built to answer it, the honest position is that the head-to-head does not exist at matched effort and harness, rather than to construct one from adjacent numbers.
What OrcaRouter routes, and what it does not
OrcaRouter is a single endpoint covering 200+ models at each provider's list rate with zero markup, which means the availability answer for this family is checkable rather than a matter of policy. Checked against our live public model list today:
• meta/muse-code — not in the catalogue. Muse Code is a harness, not a routable model, and we do not serve it
• meta/muse-spark-1.3 — not in the catalogue
• meta/muse-spark-1.2 — routable, with a 1,048,576-token context window and pricing of $1.25 per million input tokens, $4.25 per million output and $0.15 per million cached read, the same rates Meta publishes
• meta/muse-spark-1.1 — also routable, at the same three rates
So the practical shape of this on OrcaRouter is that you can reach the shipped Muse Spark generations through the same endpoint you use for everything else, and you cannot reach Muse Code or the current 1.3 generation through it. Read that as a description of today's catalogue and nothing more; availability changes and the model list is the place to check it.

Where a router earns its keep for a workload like this is in the two things a subscription cannot give you. The first is that a single API key is a routing layer that fails a request over automatically when a provider degrades, which matters when an agent run is thousands of tool calls deep and you would rather not lose it to one upstream having a bad afternoon. The second is that the request is described rather than hard-coded — you can put Muse Spark 1.2 behind the same routing DSL as any other model and change what serves it without changing your client.
Frequently asked
Is the CLI's default model the current model? No. Muse Code's default is muse-spark-1.2; the current generation is Muse Spark 1.3, announced September 2, 2026, and you reach it with --model muse-spark-1.3 or the /models command.
Which wording is right — "Muse Code 1.3" or "Muse Spark 1.3"? Muse Spark 1.3. Muse Code versions are harness releases and the current one is 1.2.1; the two numbering lines are unrelated, which is the whole reason the query "muse code 1.3" returns confused results.
If I only need audio input, which model should I configure? Muse Spark 1.2. Meta's own footnote says audio understanding in 1.3 is not fully supported and may degrade response quality, and directs audio work to 1.2 or to Muse Voice Transcribe.
Is the Contributor tier a free or trial tier? No, and it is not a discount you get automatically. It is a lower-priced variant across every Muse Spark version from 1.2 onward in exchange for permission to train on your prompts and completions; the 1.1 generation has no Contributor variant.
Bottom line
Muse Code is Meta's terminal and CI coding agent, made by Meta for its Muse Spark models, sold as three monthly plans at $5.00, $15.00 and $50.00 with no free tier, backed by pay-as-you-go Standard rates of $1.25 in and $4.25 out per million tokens and Contributor rates of $0.10 and $0.20 for the same tokens if you will let Meta train on them. The current CLI release is 1.2.1, the current model generation is Muse Spark 1.3, and the CLI still defaults to Muse Spark 1.2 — name the version and the reasoning effort whenever you quote a score, because Muse Code running Muse Spark 1.3 at max and at xhigh are not the same measurement, and the vendor's headline efficiency claim is an internal comparison rather than an independent result.
For a hands-on look at the harness itself, see our walkthrough of the Muse Code terminal coding agent.
Compared in this article3
Detected from this article · Benchmarks: Artificial Analysis · updated daily
