A generated title card headed MODEL REFERENCE with the title 'Muse Voice Transcribe' and the subtitle 'Meta's streaming speech-to-text model on Model API', above four rounded cards reading '$0.18 per hour of audio', 'Streaming WebSocket + file endpoint', '25 evaluated languages with code-switching' and 'Turn-level timestamps, not word-level', with a footer reading 'Source: Meta's own model page and docs, 2026-09-29.'
Guides & Insights

Muse Voice Transcribe: Meta's Streaming Speech-to-Text Model, Explained

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Muse Voice Transcribe is Meta's streaming speech-to-text model. It runs on Meta Model API at $0.18 per hour of audio, under the model id muse-voice-transcribe-1.0, and it is the same price whether you stream live audio or hand it a finished recording. What makes it worth a reference page rather than a one-line price check is where the work happens: speaker attribution for more than twenty voices and end-of-speech detection run inside the recognition model, not in a batch pass bolted on at the end of the stream. It is speech-to-text only — Meta's developer docs say so in as many words: it does not synthesize speech and it does not provide a speech-to-speech conversation API.

This is the page a reader lands on after searching the model's own name, so it is built to be the stable answer to two questions — what is this model, and what does an hour of audio cost — rather than an announcement. One date belongs on the record before anything else, because it is a fact about the model and not the reason the page exists: Meta Superintelligence Labs introduced Muse Voice Transcribe on 2026-09-01, in the dated announcement post Introducing Muse Voice Transcribe — the post a reader can still reach through Meta's blog index, though it no longer answers on its own URL. Everything below was read on Meta's own pages on 2026-09-29, chiefly the model page and the API's speech-to-text and overview documentation. Where Meta publishes no figure, this page says so instead of reaching for a proxy.

What it is, in Meta's terms

Meta's own documentation opens its description plainly: "Muse Voice Transcribe is Meta's speech-to-text model on Meta Model API. Transcribe live audio or an existing recording, with speech-turn detection, speaker labels, and vocabulary biasing." The overview page compresses the same thing into one sentence — streaming speaker diarization, native endpointing and voice activity detection, contextual and keyword biasing, and 25 evaluated languages with code-switching. So the output is a single transcript stream carrying three labels at once: the words, who is speaking, and whether the utterance is finished. That is the combination most production stacks currently assemble from a recognizer plus a separate voice-activity detector plus a separate diarization model, each with its own latency budget.

Meta's model page frames this as "competitive latency and streaming transcription accuracy with live attribution for 20+ speakers", and the developer docs describe the output field as "Transcript text with turn-level timing and optional speaker labels". Two things follow from that, and both matter more than the framing:

• It is an audio-in, text-out model. It is neither text-in nor text-out, and Meta is explicit that it is not a speech generator.

• It is a streaming model first. The realtime path is not a wrapper around the file path; it is a WebSocket session with its own events, modes and limits, described below.

What an hour of audio costs

Meta's model page for Muse Voice Transcribe states the price as "Build with a single model at $0.18/hour", and its specification table lists muse-voice-transcribe-1.0 at $0.18 per hour of audio and $3.00 per 1,000 minutes. Those are two unit framings of one rate, and both are printed on Meta's own page as read on 2026-09-29. The developer documentation states the same rate a third way: "Pricing is $0.18 per hour of audio processed." No figure on this page is carried over from a Meta pricing page, because there is no readable Meta pricing page to read: the docs link the rate out to a "Pricing and rate limits" page that answers This page isn't available as read on 2026-09-29, so the model page is the source of record for the rate, and it is the only one used here.

The billing detail sits in the developer docs and is worth knowing before you size a workload:

• Streaming and non-streaming cost the same. There is no premium for the realtime endpoint.

• Zero-data-retention processing is priced at parity with the Standard tier, per the docs — it is not a surcharge line.

• Billing rounds down to whole seconds, so a two-second turn bills as two seconds and not three.

• Failures that occur before a transcript is produced are not billed, and neither are 429 rate-limit responses.

• Free credits exist at the platform level, not the model level. Meta's speech-to-text documentation ends its pricing paragraph with the sentence "Platform free-tier credits apply." That is the only free-anything wording on the pages for this model, and it is deliberately non-specific: it points at a platform credit scheme rather than promising a free allowance for transcription, and no figure, duration or eligibility rule is printed for it. Read it as a credit that may reduce a bill, not as a free tier for the model. Meta's consumer-facing statement that its personal agent is "free for most of what people need" is a separate sentence about the Muse app and does not extend to Model API.

Worked example, because per-hour is hard to feel: a two-hour recorded interview is $0.36 of transcription. A thousand hours of call audio — roughly the monthly volume of a mid-size contact centre — is $180. At that scale the rate is the whole argument, and a rate is also the one thing that can move underneath you: Meta's page is the authority on it, and if the figure changes there, it changes here.

The streaming interface, endpoint by endpoint

This is the part a reference page owes a reader that a launch write-up cannot carry, so here it is in the order you meet it. Base URL for Model API is https://api.meta.ai/v1 with a Bearer token, and Muse Voice Transcribe adds two endpoints on top of it.

• Live audio — wss://api.meta.ai/v1/asr/realtime. Transport is WebSocket, and the authentication detail is a trap: you authenticate in the handshake frame, and the Authorization header is ignored. The handshake must be the first JSON text frame and must arrive within 10 seconds. A session lasts up to 60 minutes, a new WebSocket creates a new session, and there is no resume token.

• Recordings — POST https://api.meta.ai/v1/asr/transcribe. Multipart upload, and this one does authenticate with the Authorization header. Input is narrow: mono 16-bit PCM WAV at 16 or 24 kHz only. Caps are 32 MB per body and 10 minutes of audio per request.

Inside the realtime session, three modes change what you are handed. PUSH_TO_TALK is the default and does single-turn transcription. ENDPOINTING gives model-detected turn boundaries, one turn per detected speech segment, emitting speechStart, repeated transcript, speechEnd and speechComplete events — with the warning that speechEnd is not the transcript. DIARIZATION adds automatic speaker detection and attribution, emitting speaker events with labels like A and B. Partial transcripts come either cumulative, where each partial replaces the previous one, or as deltas.

Two operational details from the same documentation are easy to miss and expensive to discover in production:

• Diarization marks a possible new speaker rather than a clean speech endpoint, and Meta states it is not tuned for low-latency, voice-command use. Treat the labels as session-scoped identifiers — A in one session is not the same person as A in the next.

• Turns may overlap, so per-turn state should be keyed on the turn identifier rather than on arrival order.

• The published per-tenant ceilings are 128 concurrent streams and 16,000 streams per hour.

What runs inside the model, and what it replaces

Meta's model page makes a specific claim about architecture rather than about scores: "Speaker attribution for 20+ speakers and end-of-speech detection happen inside the recognition model" — and it explicitly contrasts that with a batch pass at the end of the audio stream. The practical difference is that there is no second stage to wait for. Speaker labels and turn boundaries arrive on the same stream as the words, so a downstream voice agent or live-captioning surface gets a completed turn and an identity without a post-processing hop.

The other capabilities Meta documents as living in the model:

• Native endpointing and voice activity detection, so you are not running your own VAD in front of the API. In the docs' phrasing, you get one completed transcript per utterance without running your own voice activity detection, and a detected end of speech does not end the session.

• Contextual and keyword biasing, with no fine-tuning. Meta's model page describes accuracy held "through contextual and keyword biasing, with no fine-tuning", which is the feature that makes streaming ASR usable on domain vocabulary — drug names, ticker symbols, product SKUs — rather than on the general-language average. The docs are precise about the limits: keywords bias recognition but do not guarantee an exact spelling, and both keywords and language bias are fixed at session start.

• 25 evaluated languages with code-switching. The docs list them: Arabic, Bengali, Dutch, English, French, German, Hebrew, Hindi, Indonesian, Italian, Japanese, Kannada, Korean, Malay, Mandarin Chinese, Marathi, Polish, Portuguese, Spanish, Tagalog, Tamil, Telugu, Thai, Turkish and Vietnamese. Code-switching — mid-sentence and mid-utterance without being told which language is coming — is the capability single-language recognizers structurally cannot offer. One caveat worth carrying: language bias is a list of languages, not free-form context, and it does not force the model to use them.

The dated announcement post goes further: it says the model was trained on 70+ languages with 25 "extensively verified" — vendor-reported, and the 25-verified list is the subset Meta stands behind. That is the coverage to plan against, not the 70.

The documented limits

A reference page is only as good as the constraints it prints, and Meta prints several.

• Turn-level timestamps, not word-level. Both the model page and the docs state it plainly: "It returns turn-level timestamps, but not word-level timestamps." Turns carry start and end times in milliseconds. If your product needs a click-to-seek per word — karaoke highlighting, per-word confidence routing, forced alignment — this is the wrong model and no amount of prompt engineering changes that.

• No confidence scores, no sound event detection, no emotion detection, no transcript reformatting. Meta lists all four as unavailable. You get words, turns, timings and speaker labels, and nothing about how sure the model is.

• No speech synthesis and no speech-to-speech conversation API. It is one direction only.

• Audio formats are narrow on the file path. Mono 16-bit PCM WAV, 16 or 24 kHz. Anything else has to be converted before it is uploaded.

Open weights, and the Muse Glimmer confusion

Muse Voice Transcribe is proprietary and served through the API — it is not an open-weight release, and Meta's documentation never claims otherwise. This matters because readers route the Muse family's open-weight path onto the wrong name. Muse Glimmer is that path: Meta's docs describe it as an open-weight multimodal model distilled from Muse Spark, distributed under a permissive Apache 2.0 licence, which you download and run on your own hardware through a runtime such as vLLM, SGLang, llama.cpp or ExecuTorch. Muse Voice Transcribe is grouped on the same page with Muse Spark, Muse Image and SAM as models you call rather than download.

So: no weights for Muse Voice Transcribe, no self-hosted option, no path to audit or fine-tune it. If a page tells you otherwise, it has conflated it with Muse Glimmer. Where a model's weights are not published, the serving platform is the whole story — and that is what the next section is about.

How a client reaches it

Meta's Model API is built to be dropped into clients you already have. The vendor's claim, printed on the API's overview page, is that Model API "is drop-in compatible with the OpenAI-compatible and Anthropic-compatible client libraries, and with OpenAI-compatible agent CLIs. Set your client's base URL, add your key, and keep the rest of your code." General to Model API, three request shapes are offered — the Responses API, the Chat Completions API and the Messages API — so an existing integration usually has a matching surface without a rewrite. One caution about scope: that compatibility sentence and those three shapes are capabilities of Model API as a whole, not lines printed on Muse Voice Transcribe's own model page. The model page's own quickstart is narrower — it says only that you "point your existing OpenAI-compatible client at Meta Model API", and that you will have a first working request in under five minutes.

One honest gap worth flagging on a reference page: Meta's documentation for this model names no official client library for Muse Voice Transcribe, and its "Client libraries" page sits under the SAM section rather than the Voice one. What the speech-to-text guide gives you instead is a concrete dependency list — Python 3.9 or later with the websockets and sounddevice packages — plus a voice-API fundamentals cookbook. So the drop-in claim describes Model API's request surface in general; the realtime ASR endpoint itself is a WebSocket with its own handshake rules and no documented wrapper.

A screenshot of Meta's model page for Muse Voice Transcribe (captured 2026-09-29), showing the headline 'Competitive latency and streaming transcription accuracy with live attribution for 20+ speakers. Build with a single model at $0.18/hour.', a 'Meet Muse Voice Transcribe' panel of capability cards headed 'Competitive transcription accuracy', 'Live speaker and turn awareness' and 'Production pricing', a 'Muse Voice Transcribe benchmarks' section labelled 'AA-WER Streaming Index - Final Transcription Diarization' with 3.9% and 3.9%, and the pricing row reading muse-voice-transcribe-1.0 at $3.00 per 1,000 minutes and $0.18 per hour.

What Meta has not published

An absent benchmark is a fact about the documentation, so here it is stated as one. Meta's model page for Muse Voice Transcribe carries no printed accuracy table: its accuracy evidence is delivered as three image assets — a streaming word-error-rate leaderboard, a diarization error-rate comparison and a streaming accuracy index — with no figures in the extractable text. The model page itself gives no word error rate, no diarization error rate, and no language-by-language breakdown.

The announcement post does carry vendor-reported figures, including a streaming word error rate and an average diarization error rate, and those are the numbers we wrote up when the model launched; they are Meta's own measurements and remain unreproduced by a third party. This page does not re-run that comparison, because re-running a release-day vendor benchmark on a reference page would dress an unreproduced figure up as a settled one. What can be said plainly is narrower: Meta publishes no head-to-head evaluation against a named rival at matched effort and matched harness for this model, so any comparison you read anywhere is a comparison of vendor numbers collected under different conditions.

One date also stays unestablished on purpose. The model page carries no release date at all — its only version marker is the -1.0 in the model id. The 2026-09-01 date in this piece comes from that dated announcement post, and it is used here to date the model rather than to report an event.

A generated reference scoreboard titled 'Muse Voice Transcribe — the reference', listing six rows for the model: price $0.18 per hour of audio, interface as a realtime WebSocket at wss://api.meta.ai/v1/asr/realtime plus a file endpoint, speaker attribution for 20+ speakers inside the recognition model, turn-level timestamps without word-level times, 25 evaluated languages with code-switching, and proprietary weights with no open download.A screenshot of the Speech to text page in Meta's Model API documentation (captured 2026-09-29), showing the opening sentence 'Muse Voice Transcribe is Meta's speech-to-text model on Meta Model API', an 'Available endpoints' table contrasting wss://api.meta.ai/v1/asr/realtime for live audio against POST https://api.meta.ai/v1/asr/transcribe for a recording and noting that the Authorization header is ignored on the WebSocket, and a Features table whose rows include language biasing across 25 supported languages with code-switching, vocabulary biasing, speech-turn detection, speaker labels, partial transcripts and progress events.

Who should reach for it, and who should not

The case for Muse Voice Transcribe is narrow and strong: live audio where you need words, speaker identity and turn ends on the same stream, at a price that is low enough to run continuously rather than on demand. Voice agents, live captioning, meeting and call intelligence, dictation and high-volume transcription are the workloads Meta names, and they are the workloads the interface is actually shaped for — a 60-minute WebSocket with endpointing modes and session-scoped speaker labels.

The case against it is equally specific. If you need word-level timestamps, per-word confidence, sound-event or emotion detection, or transcript reformatting, this model does not have them, and that is a documented absence rather than a bug. If your audio arrives in formats other than mono 16-bit WAV at 16 or 24 kHz and you are using the file endpoint, you pay a conversion step you would not pay elsewhere. And if your requirement is batch transcription of archived recordings where accuracy is the only axis, the streaming pitch buys you nothing — the price is the same on both paths, so you are paying for real-time capability you will not use.

The one thing to check before you commit a production path is the figure this page exists to answer, and it is the one figure that can change without the model changing at all: $0.18 per hour of audio, as printed on Meta's own model page today. Meta's page is the authority on it. When it moves, everything sized on it — the thousand-hour month, the per-interview cost, the build-versus-buy arithmetic — moves with it.