Hero title card for MAI-Image-2.6, Microsoft's #2 text-to-image model on Arena, showing an Arena score of 1,336, a No. 2 rank, and a +79 Elo gain over MAI-Image-2.5.
Guides & Insights

MAI-Image-2.6: Microsoft's New #2 Image Model Debuts on Arena Before Its API Is Live

Author

Jim Song

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Microsoft's MAI-Image-2.6 just became the second-highest-rated text-to-image model on the Arena leaderboard — while still having no public API. Announced on August 10, 2026, the model sits at No. 2 with an Arena score of 1,336, trailing only GPT Image 2 Medium (1,381) and ahead of Grok Imagine Image 2.0 (1,316), Meta's Muse Image, and Nano Banana Pro. Its predecessor, MAI-Image-2.5, ranks tenth.

The striking part is the context, not just the rank. The entry that topped out at No. 2 is literally named "mai-image-2.6-preview" and carried roughly 3,500 votes in that snapshot. It climbed eight places and gained +79 Elo over MAI-Image-2.5 in a single step, and Microsoft is promising API and Foundry access "soon," with early reporting pointing to next week. That makes this the strongest independent signal yet that Microsoft's in-house image effort has closed the gap with OpenAI — and most of you still can't call it through an API.

The key numbers

Arena text-to-image rank: No. 2 — score 1,336±11, preliminary, ~3,488 votes

Gap to No. 1 (GPT Image 2 Medium): 45 Elo

Lead over No. 3 (Grok Imagine Image 2.0): 20 Elo

Jump over MAI-Image-2.5: +79 Elo overall, +91 Elo on text rendering

Predecessor's position: MAI-Image-2.5 at No. 10 (1,256)

Availability: Arena now; MAI Playground later this week; APIs/Foundry expected next week

Scoreboard for MAI-Image-2.6: Arena T2I rank No. 2 (1,336 Elo, preliminary), gap to GPT Image 2 Medium -45 Elo, gain vs MAI-Image-2.5 +79 Elo, text rendering +91 Elo, best category 3D and modeling No. 1, API status expected next week. Footer notes Arena figures as of Aug 10, 2026 and that vendor claims are unreproduced.

What MAI-Image-2.6 is

MAI-Image-2.6 is the newest model in Microsoft's in-house image generation line, built under Mustafa Suleyman's Microsoft AI. The family has moved fast: MAI-Image-1 was the foundation, MAI-Image-2 (March 2026) delivered the photorealism and text-rendering jump, MAI-Image-2.5 (June 2026) added professional-grade output and editing and became the engine inside Bing Image Creator, PowerPoint, and OneDrive, MAI-Image-2.5-Pro (July 2026) added an 8K premium tier — and 2.6 now lands roughly six weeks after its predecessor.

The cadence matters as much as the model. Microsoft is shipping image improvements on a six-to-eight-week cycle, closer to how frontier labs ship text models than to how the image market has historically moved. A No. 2 Arena debut is what that cadence looks like once it compounds.

The Arena result, read carefully

Arena (the rebranded LMArena) ranks image models by blind human preference: voters see two images generated from the same prompt and pick the stronger one. That makes it an independent measure of what people actually prefer rather than a vendor-run benchmark — which is why a No. 2 debut carries real weight. It's also why the number needs caveats: the 1,336 score is preliminary, based on about 3,500 votes versus more than 70,000 for the model in first place, and the entry is labeled "preview." Early-vote Elo moves; treat the direction as solid and the exact score as provisional.

The Arena text-to-image leaderboard as of August 10, 2026, showing MAI-Image-2.6 in second place at 1,336 Elo with 3,488 votes, behind gpt-image-2 (medium) at 1,381 and ahead of grok-imagine-image-2.0 at 1,316.

Where the gains are concentrated is more interesting than the headline number. Arena ranked MAI-Image-2.6 first in 3D imaging and modeling (up from sixth) and second in cartoon/anime/fantasy (up from eighth), product and branding and commercial design (up from seventh), text rendering (up from eighth, with the +91 Elo gain), and art (up from fourth). Microsoft says it improved in every measured category. That spread — commercial design, 3D, text, and art — is the signature of a general-purpose image model rather than a one-trick specialty.

What's actually new (vendor-reported)

On capabilities, Microsoft's own descriptions for MAI-Image-2.6 are:

• Stronger text rendering, portraits, and 3D imagery

• More polished commercial and photorealistic output across product, branding, and cinematic use cases

• Work across multiple reference images in a single generation

• Richer grounding — better adherence to the specifics in the prompt

• Greater control over reasoning, format, and resolution

Every one of those is vendor-reported and not yet independently reproduced. The only independent evidence so far is the Arena jump — which is genuinely independent but preliminary. Until the API opens and third parties can run controlled evals, treat the capability list as Microsoft's claim and the leaderboard as the working signal.

When you can actually use it

The model is live to test today, but only through Arena. Microsoft says the MAI Playground gets it later this week, and API access plus Microsoft Foundry are "rolling out soon" — early reporting points to next week. The practical tell to watch: when MAI-Image-2.5 became the default in Bing Image Creator and rolled into PowerPoint and OneDrive, it meant Microsoft considered it production-safe. A similar 2.6 rollout into those surfaces will be the signal that the preview label is coming off.

Microsoft's announcement that MAI-Image-2.6 launches at No. 2 on the Arena leaderboard ahead of Google, Meta and xAI, with category highlights for text rendering, 3D imaging and modeling, branding and commercial, portraits, and photorealistic and cinematic output.

What it might cost

No pricing for 2.6 has been published. The anchor is the rest of the family, all vendor-reported:

• MAI-Image-2.5: $5 per million text-input tokens, $8 image input, $47 image output

• MAI-Image-2.5-Flash: $1.75 / $1.75 / $19.50 — the high-volume tier

• MAI-Image-2.5-Pro (July 2026): $5 / $8 / $106 — 8K output, 96.8% text-rendering accuracy

That is a five-fold spread between the cheapest and priciest image-output rates Microsoft already sells, so 2.6's production cost will hinge on which tier it slots into. At 2.5's $47-per-million output rate, an image consuming roughly a thousand output tokens lands near five cents — but Microsoft does not publish per-image token counts, so treat that as arithmetic, not a quote.

Pricing is also where a routing layer earns its keep once the API opens. On OrcaRouter, whatever Microsoft charges is passed through at 0% markup — provider list price, nothing added — so a launch discount or a later price cut lands on your side of the bill the same day it's announced. No renegotiating and no second contract; the model shows up on the same key as everything else.

Where this leaves the image-model race

The top of Arena's text-to-image board right now: GPT Image 2 (Medium) at 1,381, MAI-Image-2.6 at 1,336, Grok Imagine Image 2.0 at 1,316, Reve 2.1 at 1,302, and Meta's Muse Image at 1,282. Five models within roughly 100 Elo at the top. The quality gap that used to separate the leader from the pack has compressed to the point where "best image model" is no longer a useful question — editing, grounding, speed, and price are.

Microsoft's bet is broader than any single leaderboard crown: an in-house model stack spanning image, voice, code, and reasoning, distributed through Copilot, Bing, and Office. The Arena rank is the proof-of-work; the distribution is the actual product.

How to try it without betting production on it

The Arena preview is free — go generate against it today if you want the raw signal. Once the API opens, the sensible way to adopt a brand-new model with a "preliminary" score is to route a slice of traffic to it and keep your proven model as the fallback. That is exactly the failure mode a routing layer is built for: on OrcaRouter you would point one endpoint at MAI-Image-2.6 once it is hosted on an upstream provider and let automatic failover catch the edge cases a 3,000-vote leaderboard rating cannot see. One API key, 200+ models, provider prices passed through.

What to watch next

• Whether MAI-Image-2.6 shows up on the image-editing leaderboard — MAI-Image-2.5 is already a strong editor, and 2.6's editing scores are not out yet

• The API price and which tier it lands in

• Whether Bing Image Creator and Copilot switch their default from 2.5 to 2.6

• Independent benchmarks once the API opens

• Whether the "preview" label drops — the signal that Microsoft is confident enough to sell it

FAQ

When is MAI-Image-2.6 available through an API?

Not yet. It is on Arena now as a preview build; Microsoft says API and Foundry access are "rolling out soon," with early reporting pointing to next week. No firm date or price has been published.

How close is it to GPT Image 2 Medium?

45 Elo on the current Arena snapshot — close enough that the No. 2 / No. 1 difference sits within what vote flow can move, and the top five models are all within roughly 100 Elo.

Is the No. 2 Arena rank reliable?

It is an independent, blind human-preference result, which is exactly what makes it meaningful — but it is preliminary (about 3,500 votes) and attached to a preview build. Trust the direction, not the decimals.

Bottom line: MAI-Image-2.6 is the first Microsoft image model that genuinely belongs in the frontier conversation, and it arrived before the API did. If the price lands anywhere near the Flash tier, it becomes one of the most interesting value plays in image generation this year. The API cannot open soon enough — and when it does, the price and the editing scores will tell you whether No. 2 was a debut or a plateau.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

Contact us

Join our community

DiscordEmailXGitHubYouTube