
Meta Muse Image: The Model, the API, and the Flat $0.01 Image
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 1009 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1189 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 22 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 108 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
If you searched the disambiguated name and landed here, the short answer is this: Meta Muse Image is Meta's agentic image generation and editing model, it is proprietary rather than open weights, and you call it through Meta's own Model API under the identifier muse-image-1.0 at a flat $0.01 per generated image, with search grounding included in that price. That is the whole answer to the question the search is asking, and everything below is the reference for it — what the model does, how to reach it, what the unit of billing actually is, what Meta documents about provenance, and where the documentation stops.
Two things about this page's warrant are worth stating plainly, because this family has already been announced more than once and a second announcement piece would be a duplicate of coverage we already run. This is the canonical reference page for the name: the destination a reader reaches after searching it. Its warrant is standing demand that we measured first-party on September 29, 2026 — readers type the disambiguated form of this name across three spellings, in the hundreds of impressions, at a position we hold by accident of other pages rather than by having written a page shaped like this one, and it converts at close to nothing because nothing we publish is shaped like a reference. That measurement is a fact about what people type and nothing else; it says nothing about the model's quality or popularity in the world. And the datable anchor is used here as a fact to date the model, not as an event to report: Meta's launch post for Muse Image and Muse Video carries a dateline of July 7, 2026 and describes the model as "available today" across Meta's consumer surfaces on that date, so July 7, 2026 is the launch date this page works from. Every figure below was read from Meta's own pages on September 29, 2026 unless the sentence says otherwise.
What Muse Image actually is
Meta's own one-line description is the place to start, because it is doing real work rather than marketing: Muse Image is "an agentic image generation, grounded by search and priced for production volumes at $0.01/image." On Meta's model page the capability surface is broken into four working modes — image generation, precise editing, anchored composition, and multi-turn refinement — and the developer documentation compresses the same thing into a sentence: "Muse Image generates and edits images from a conversation."
The distinguishing claim, in Meta's words, is that Muse Image "is an image model that reasons before it renders." Meta's launch post makes that concrete rather than notional: "Instead of directly mapping prompts to images, Muse Image operates as an agent: it invokes search and coding tools to improve accuracy, self-refines its own generations, and improves through scaling test-time compute." In practice that means the model plans before it draws, pulls live references from the web when a prompt touches real-world facts, and writes and runs code when precision matters — the path Meta describes for accurate plots and QR codes — then checks the result before returning it.
The three capabilities Meta names first are the ones to evaluate it on, and they are stated in Meta's own sentence: Muse Image "follows instructions faithfully, edits with precision, and composes from multiple references." The editing claim is the sharpest of the three, because it is a claim about restraint rather than about output: Meta's model page says the model "is designed to change exactly what you specify and leave everything else untouched," and supports iterative refinement across turns rather than one-shot edits. Composition is where the "anchored composition" mode lives — anchoring "a whole series of generations on a small set of reference images so the character, style, and setting stay consistent from one image to the next," which is the capability that matters if you are producing a catalogue or an ad set instead of a single picture. Taken together these are the claims that separate an image model you generate with from one you build a pipeline on, and they are vendor claims: Meta publishes no independent benchmark of instruction-following or edit fidelity for this model.
The agentic framing is not only about tools. Meta documents Muse Image as pairing with Muse Spark, its reasoning model, so that "the two models to share tools and plan jointly for powerful agentic media generation" — the documented example being multi-part output like animated GIFs, pages with embedded images, and small interactive games. If you are building anything that mixes generated imagery with generated text or code, that pairing is the part of the design worth understanding, and it is the reason the model is served on the same API and the same auth as Muse Spark rather than on a separate image endpoint with its own credentials.
How you call it: the API surface
Muse Image is reached through Meta's Model API, and the developer documentation is specific about the shape of that call. The model identifier is muse-image-1.0, passed as the model field on every request. Requests go to the base URL https://api.meta.ai/v1 with a bearer token — the documentation notes that it uses the same base URL and auth as Muse Spark — and there are two image endpoints plus a conversational one:
• Generate — POST /v1/images/generations, text prompt in, image out.
• Edit — POST /v1/images/edits, prompt plus input images, for image-to-image work.
• Refine across turns — POST /v1/responses, the Responses API, which is how the documented multi-turn flow works.
On the images endpoints the documented parameters are few and worth committing to memory, because they are where the cost and the output format are decided. n takes 1 to 10 and defaults to 1. size takes a "WxH" string such as 1792x1024, and the documentation is explicit that on these endpoints it "sets aspect ratio only, not exact pixels" — do not read it as a resolution guarantee. response_format is b64_json by default or url, and output_format is webp by default with png and jpeg available. The agentic behaviour is switchable: the images endpoints take a tool_enablement extension covering enable_image_search, enable_web_search and enable_shell, and a reasoning_strength of high (the default) or low. Turning reasoning down is the lever for cost and latency, and it is the first thing worth A/B testing on your own prompts rather than taking on faith.
On the Responses API the same controls live inside an image_generation tool object in tools, with image search, web search and shell all enabled by default, and the output array comes back in a fixed order: a reasoning item summarising what the model planned and looked up, a message item, then one image_generation_call item per image. The image itself is in that item's result field as base64, and its id is documented as a signed handle you pass back to keep editing the same image in a later turn. That handle is the mechanism behind multi-turn refinement — it is how you avoid re-uploading an image to change one detail of it. Both images endpoints also accept stream: true, emitting image_generation.completed for generation and image_edit.completed for edits. Model API is documented as drop-in compatible with the OpenAI SDK, the Anthropic SDK and OpenAI-compatible agent CLIs, so an existing image call is a model-string change rather than a rewrite.
One caution that belongs in the reference rather than in a footnote: Meta has not published rate limits for Muse Image that we could read. The rate-limits page on Meta's own developer site returned a shell with no figures on it, the same as the pricing page. If throughput matters to your design, that is an unknown you will discover from the API rather than from the documentation.

The price, and the unit it is priced in
One cent per image. That figure is read from Meta's own model page for Muse Image, which lists the model in a pricing table with search grounding and reasoning both marked as included and the price given as $0.01/image; the developer documentation restates it as "Muse Image is billed at a flat $0.01 per generated image." The unit is the image, not the token — the documentation says so directly ("Muse Image isn't priced per token") — and that matters because image APIs that report usage in tokens are usually billed in tokens. Muse Image returns a usage block with input and output token counts, and those counts are informational only.
Three billing details follow from the unit. First, n is a multiplier: a request that returns ten images bills for ten images, so the 1–10 parameter is a spend control as much as a convenience. Second, failure is not billed — images that fail to generate or that are removed by safety filtering before they are returned are not counted. Third, search grounding is inside the per-image price rather than a separate line item; Meta's documentation states it is "part of the per-image price, so it carries no separate search-grounding charge." That last point is the one that makes the rate meaningful for agentic work: a model that searches the web before rendering can generate several tool calls behind a single output, and Meta is stating that the tool calls are not billed separately from the image.
We could not corroborate the number from a rate card, and that is worth saying rather than skipping. Every pricing URL under Meta's developer domain that we fetched returned a page shell with no dollar figures in it, so the model page and the image-generation documentation are the two places this figure exists as a readable statement. If you need a second source for procurement, there is not one on Meta's site today.

What one cent buys you in practice is a decision change rather than a discount. Meta's own framing on the model page is that the model "lets you run workloads that frontier pricing previously put out of reach" — ad variant generation, catalogue imagery and per-user personalisation — and at this rate a hundred thousand images list at roughly $1,000. If you are weighing that rate against the premium image models you may already pay for, the useful move is to run both on your own prompts and your own evaluation, and one cent per image is what makes that comparison cheap enough to actually do.
Provenance: what Content Seal does and does not settle
Meta ships an invisible watermarking system with this model, and documents it precisely enough to be worth quoting: "Muse Image includes Content Seal, our invisible watermarking system. Images created by Muse Image in the Meta AI app and on meta.ai carry a hidden provenance signal that stays intact — even when cropped, compressed, resized, or screenshotted." Meta says it plans to extend Content Seal to video, and previews a detection tool that lets you check whether an image carries the mark.
Read the scope of that sentence carefully, because it is narrower than it first appears. The provenance signal is described as being carried by images created in the Meta AI app and on meta.ai — the consumer surfaces — and Meta's model page and API documentation say nothing about whether images generated through the Model API carry the same mark. We could not find a statement either way, so treat Content Seal as a consumer-surface provenance feature and do not assume it travels with an API-generated file until you have tested output from your own account. Independently, the one test of the detector we have reported came from Reuters, which found it failed to verify a majority of watermarked images after cropping; that finding and its figure are covered on our explainer of this model, which is the page to read for the provenance argument rather than this one.
Where it runs
Meta's own availability statement, from the launch post: Muse Image "is available today across the Meta AI app and on meta.ai, Instagram Stories in the US, and WhatsApp in limited countries, and is coming soon to Facebook." Those are Meta's words and the surfaces are Meta's own products — there is no waitlist, no invite code and no separate signup for any of them. On the developer side the model is served through Model API, which is what the identifier and endpoints above are for. Meta's model page also lists third-party inference platforms among the ways to run it; we are not naming them here, and nothing in this page should be read as a claim about terms or rates on any surface other than Meta's own API.
Meta has not published a date for when Muse Image landed on Model API, and Meta's developer documentation states none. The one date this page works from is the launch date of the model itself, July 7, 2026, established from the dateline of Meta's own announcement post. If you are dating the API surface rather than the model, say what Meta says, which is nothing.
Open weights or proprietary? Proprietary — and the family has both
This is the question readers most often get wrong in this family, so it is worth answering directly. Muse Image is proprietary. Meta's developer documentation draws the line explicitly: the models you call over Model API — Muse Spark, Muse Image, Muse Voice Transcribe and SAM — are served, and "Muse Glimmer takes a different path: you download the open weights and run it on your own hardware instead of calling it through Model API," under a permissive Apache 2.0 licence. Muse Glimmer is the open-weight path in this family, and it is a different model: a multimodal model distilled from Muse Spark, not an open release of Muse Image. Nothing on Meta's page for Muse Image states a licence or offers weights, because there are none to offer. If self-hosting is a hard requirement, Muse Glimmer is the model to evaluate; if you want the image model described here, you call Meta's API and you pay per image.

What Meta claims it ranks, labelled as Meta's claim
Meta's launch post states: "Muse Image holds the No. 2 spot on Arena for text-to-image, single-image editing, and multi-image editing as measured by human preference Elo rankings at the time of writing." Two things about that sentence matter and both come from Meta's own wording. It is a vendor-reported ranking, not an independent score, and it is an human-preference Elo standing rather than a benchmark result on a fixed harness — a distinction that makes a real difference when the gap between the top two positions is smaller than the margin of an Elo rating built on a limited number of votes. The charts Meta publishes alongside the claim are dated "as of July 5, 2026," and the claim itself is qualified "at the time of writing," so it is a snapshot of a live leaderboard quoted by a party with an interest in it, not a settled standing.
The same post includes a scaling figure that Meta labels for you: the chart showing quality rising with reasoning is captioned "Elo from internal ablation." That is a vendor-internal measurement, and it is not comparable to a third-party board. Meta's argument is that quality improves with test-time compute in an approximately log-linear relationship and that reasoning and tool use compound when combined — an internal ablation, stated as one.
Independent positions for this model exist, on boards that are not Meta's, and we have reported them: our explainer carries the Artificial Analysis image boards and the arena positions with their figures, and the two matchup pages put those numbers beside named rivals. If you want the independent picture rather than the vendor's, that is where to read it. This page does not reprint those figures, and one of them has to be read carefully anyway, because Artificial Analysis only ranks image models that ship a callable API — so an independent position on Muse Image is a statement about Model API, not about the Meta AI app.
What the documentation does not answer
A reference page is only useful if it says where the record stops, so here is what we could not establish from Meta on September 29, 2026. There is no published rate limit for Muse Image. There is no published latency profile, no success rate and no independent evaluation of the agentic loop under production traffic — the per-image price is documented but the operational envelope is not. There is no benchmark of instruction-following or edit fidelity from any independent party, and Meta publishes no such number of its own; the benchmark graphic on Meta's model page carries no readable figures. There is no licence statement, because the model is not licensed for download. And the date Muse Image became callable on Model API is not stated anywhere in Meta's documentation. None of those absences is a reason to avoid the model — a one-cent image with included search grounding is worth testing on that basis alone — but each is a reason to measure rather than assume before a pipeline depends on it.
The verdict for a developer
Meta Muse Image is worth evaluating if your problem is volume: catalogue imagery, ad variants, per-user personalisation, anything where the frontier per-image rate was the thing that killed the project. The unit price is documented, the unit is unambiguous, failures are not billed, and search grounding is inside the rate rather than billed alongside it, which is unusual enough in this market to be the headline. It is not worth adopting blind, because everything beyond the price — throughput, latency, how reliably the agentic loop closes on your prompts, whether provenance signalling travels with API output — is undocumented, and the vendor's own ranking claim is a human-preference standing quoted at a moment in time rather than a measured result. The honest shape of a decision here is a one-cent experiment against the image model you already ship with, on your own prompts, with the cheap failure mode and a proven fallback behind it. For what this model is, how it is called and what it costs, this page is the reference; for whether it should be in your stack, the numbers you generate yourself are the only ones that count.
