A generated hero card titled 'Claude Haiku 5.5 vs MiniMax M3.1 Flash Preview' with the kicker 'ONE HAS A PRICE. THE OTHER HAS A TOKEN PLAN' and chips reading $0.10 / $0.50 per 1M against not published, 1M context each, and AA Index 43 against none published.
Guides & Insights

Claude Haiku 5.5 vs MiniMax M3.1 Flash Preview: One Has a Price, the Other Has a Token Plan

Author

Alistair Wren

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MiniMax M3.1 Flash Preview has no published per-token price. Not a high one, not a promotional one — none. It is available only through the vendor's Token Plan and its Code surface, and the developer documentation says so in as many words. Claude Haiku 5.5, released October 7, 2026, costs $0.10 per million input tokens and $0.50 per million output tokens up to a 100,000-token prompt. So the first thing to say about this matchup is that it is not a price comparison, because only one of the two models has a price. What it is instead is a comparison between a metered API and a subscription — and that difference decides more about your costs than any benchmark will.

The second thing to say is that MiniMax M3.1 Flash Preview has no model card, no published benchmarks, no weights and no independent score that we can find. It appeared on September 27, 2026 and it is, by the vendor's own labelling, a preview reached through a subscription rather than an API. That is not a criticism of the model; it is a description of what you can and cannot know about it. Everything below keeps that line visible.

What each one actually is

Claude Haiku 5.5 — Anthropic, ID claude-haiku-5-5, released October 7, 2026. Text and image in, text out. A 1M-token context window, 128,000-token maximum output (300,000 via a beta header on the Batch API), June 2026 training cutoff, retirement not before October 7, 2027. Adjustable effort across Low, Medium, High, Xhigh and Max with Medium as the default and adaptive thinking on by default. Metered per token, with cache reads at $0.01 per million and Batch at half price.

MiniMax M3.1 Flash Preview — MiniMax, live since September 27, 2026. A 1M-token context window. Text, image and video in, text out. Thinking is always on, with an effort control running from low to max. The API is Anthropic-compatible, OpenAI-compatible and OpenAI-Responses-compatible, which makes it a near drop-in for either SDK. Access is through the Token Plan and MiniMax Code only; there is no pay-as-you-go path and no published rate.

A capture of the MiniMax developer documentation text-model page containing the line stating that MiniMax-M3.1-Flash-Preview is available only through Token Plan and MiniMax Code for now, alongside the subscription-key instructions and the supported API endpoint formats.

• The comparison, dimension by dimension

• Price — Claude Haiku 5.5 $0.10 per million input and $0.50 output to 100K tokens, then $0.50 and $2.50; MiniMax M3.1 Flash Preview has no published per-token rate, only a Token Plan subscription.

• Context window — 1,000,000 tokens each. This is the one dimension where they are genuinely level.

• Maximum output — 128,000 tokens for Claude Haiku 5.5, with a 300,000-token beta path on Batch; not published for MiniMax M3.1 Flash Preview.

• Input modalities — text and image against text, image and video. The video input is a real differentiator and MiniMax wins it.

• Effort control — five positions, default Medium, on Claude Haiku 5.5; low to max with thinking always on for MiniMax M3.1 Flash Preview.

• Independent score — Artificial Analysis Intelligence Index 43, Max configuration 43.40, second of 182 for Claude Haiku 5.5; none published for MiniMax M3.1 Flash Preview.

• Cost per finished task — $0.21 per Artificial Analysis index task for Claude Haiku 5.5; not computable for MiniMax M3.1 Flash Preview without both a score and a rate.

• Output speed — 243.4 tokens per second for Claude Haiku 5.5; unreported for MiniMax M3.1 Flash Preview.

• Open weights — neither. Claude Haiku 5.5 is proprietary; MiniMax has not released weights for M3.1 Flash Preview.

Why a subscription is a worse fit than it sounds

There is a version of this matchup where the Token Plan is the better buy, and it is worth being fair to it. If you are a person using MiniMax Code interactively, a subscription caps your cost at a known monthly figure regardless of how much you use, and per-token pricing has no equivalent ceiling. For a heavy interactive user, a flat plan beats a meter almost every time.

That logic inverts the moment the workload is programmatic. A subscription is priced for a human's usage pattern — bursty, interactive, self-limiting. An agent, a batch job or a production endpoint has no such pattern: it runs as much as the queue demands. Putting that traffic on a plan priced for interactive use means either hitting a limit you cannot see, or paying for a seat when you needed throughput. And the second-order problem is worse: with no published per-token rate, there is no way to forecast cost as volume grows, no way to compare against any other model on price, and no way to decide whether a cheaper model would have done the job. A price you cannot compute is a budget you cannot defend.

Claude Haiku 5.5 is the opposite problem in the best way. The rate is published, the tiers are published, the cache discount is published, and the only trap is structural rather than hidden: cross 100,000 tokens of prompt and both rates multiply by five. That is a cliff you can design around. An unpublished rate is a fog you cannot.

The 1M-token tie, and why it does not mean what it looks like

Both models advertise a 1M-token context window, and that is the headline both vendors lead with. It is also the dimension where the tie is most misleading, because context windows are not the same thing as usable context. Claude Haiku 5.5 publishes a 128,000-token maximum output alongside its window and charges five times the base rate above 100,000 input tokens — a pricing design that tells you plainly that long prompts are allowed but not the intended default. MiniMax publishes no maximum output for M3.1 Flash Preview and no long-context pricing at all, because it publishes no pricing at all.

Long context is exactly where the cost difference between the two architectures would show up most clearly, and it is exactly the place where one of them gives you no number. If your workload is genuinely long-context, that absence is the deciding fact, not the shared 1M figure.

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs MiniMax M3.1 Flash Preview — the scoreboard'. Claude Haiku 5.5 reads price $0.10 / $0.50 per million, 1M-token context, 128K max output, text and image input, AA Index 43, billed per token; MiniMax M3.1 Flash Preview reads price not published, 1M-token context, max output not published, text image and video input, no published independent score, billed by subscription plan. A footer notes the Claude figures are per Artificial Analysis and the MiniMax figures are vendor-reported with no independent evaluation.

If what you actually want is MiniMax M3

MiniMax M3.1 Flash Preview is a preview, gated behind a plan, and unscored. MiniMax M3 is none of those things. It is the open-weight flagship MiniMax shipped on May 31, 2026: a 1M-token context window, native multimodality taking text, images and video in and returning text, built on MiniMax Sparse Attention, and priced at $0.30 per million input and $1.20 per million output. Artificial Analysis scores it at Intelligence Index 29 with a cost per index task of $0.51, 91.7 tokens per second of output and a 2.16-second time to first token, and an 80% cache discount. It is on OrcaRouter's catalogue today as minimax/minimax-m3.

Which makes a concrete point about this matchup that a spec sheet hides. The MiniMax model you can actually buy per token is a generation behind the preview, scores 29 on the independent index against Claude Haiku 5.5's 43, costs about twice as much per finished task at $0.51 against $0.21, and runs at about a third the output speed. The preview might well beat all of that. Nobody outside MiniMax knows, and there is no way to find out from a subscription page.

A capture of the OrcaRouter model page for minimax/minimax-m3 showing the MiniMax M3 listing with its Vision, Tools, JSON and Reasoning capability badges, the MiniMax byline and May 31, 2026 date, the 1M-token context description, the $0.30 and $1.20 per-million pricing and the Get the MiniMax M3 API call to action.

Calling these two from one place

MiniMax M3.1 Flash Preview is available through MiniMax's Token Plan and MiniMax Code. Claude Haiku 5.5 is available through Anthropic's own API and, from there, wherever Anthropic models are resold. Neither is on OrcaRouter's catalogue — what we route from this pairing is minimax/minimax-m3, alongside anthropic/claude-haiku-4.5, anthropic/claude-sonnet-5.5, anthropic/claude-opus-5.5 and anthropic/claude-fable-5.1, and we will not imply we carry the two models this article is about.

What the routing layer genuinely offers here is an escape from the two-billing-shape problem. minimax/minimax-m3 is metered and on our catalogue; Claude Haiku 5.5 is metered and on Anthropic's. Both speak an OpenAI-compatible interface, so the A/B between them is a model string rather than a rewrite, and OrcaRouter passes provider list price through at 0% markup, meaning a MiniMax price change reaches your bill the day it happens. The Token Plan model is the one piece that cannot be normalised that way — a subscription has no per-token price to pass through, which is the same problem stated from the billing side.

How to decide without a score

If you need video input, MiniMax M3.1 Flash Preview is the only one of these two that takes it, and if your workload is interactive rather than programmatic, the Token Plan is a reasonable thing to buy. If you need a number you can put in a budget, a score you can check and a window you can price, Claude Haiku 5.5 is the only one of the two that offers any of that. And if you want the MiniMax family on a metered API with an independent score attached, the honest recommendation is to use the generation that is actually purchasable per token rather than waiting on a preview whose terms have not been published.

Both speak an OpenAI-compatible interface, so the A/B between them is a model string rather than a rewrite, and OrcaRouter passes provider list price through at 0% markup , meaning a MiniMax price change reaches your bill the day it happens.