Hero card for MAI-Image-2.5-Pro, Microsoft's quality-focused image model, showing its #1 rank in image editing and #7 in text-to-image on the Artificial Analysis leaderboard
Guides & Insights

MAI-Image-2.5-Pro Takes the #1 Spot in Image Editing: What Microsoft's Independent Win Actually Proves

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

MAI-Image-2.5-Pro is now the top-rated image-editing model on Artificial Analysis's leaderboard — an independent, blind-vote board — while still ranking only seventh at plain text-to-image. That split is the story: a model that has been in public preview for under a month, built entirely in-house by Microsoft AI, has overtaken GPT Image 2, Reve 2.1, and its own sibling MAI-Image-2.5 on the editing board, yet trails most of the same field when the prompt is "draw me something from nothing." Read both numbers together and you get a precise picture of what Microsoft shipped: a specialist in changing existing images, not a general-purpose generator.

The two numbers, side by side

As of the August 2026 snapshot, Artificial Analysis's Image Editing Arena puts MAI-Image-2.5-Pro at No. 1 with Elo 1,272, built on roughly 6,100 blind user comparisons. The chasing pack is tight: Reve 2.1 at 1,262, then MAI-Image-2.5 and GPT Image 2 (high) tied at 1,256, GPT Image 1.5 (high) at 1,250, Qwen-Image-3.0-Pro at 1,249, and Go​gle's Nano Banana 2 at 1,249. On the Text-to-Image Arena it sits at No. 7 with Elo 1,292, behind GPT Image 2 (high) at 1,369, Reve 2.1 at 1,322, Nano Banana 2 at 1,320, GPT Image 1.5 (high) at 1,310, MAI-Image-2.5 at 1,304, and Nano Banana Pro at 1,296.

Editing: No. 1 — 1,272 Elo, ~6,100 samples

Text-to-image: No. 7 — 1,292 Elo, ~11,500 samples

Gap to No. 2 on editing: 10 Elo over Reve 2.1

Gap to No. 1 on text-to-image: 77 Elo behind GPT Image 2 (high)

What makes the editing rank worth more than a vanity number is the mechanism. The Artificial Analysis arena is blind human preference: voters see the same input image and instruction, get two edited results, and pick the stronger one. No vendor script writes the score. So a #1 there is the closest thing the market has to "people liked this editor's output most" — and it is a recent climb, not a launch-day blip. Treat the direction as solid and the exact Elo as provisional; low-vote boards move.

Screenshot of the Artificial Analysis Image Editing Leaderboard showing MAI-Image-2.5-Pro at No. 1 with Elo 1,272 and 6,133 samples, ahead of Reve 2.1 at 1,262, MAI-Image-2.5 at 1,256, and GPT Image 2 (high) at 1,256

The text-to-image board, captured the same day, is the same model in a different light — no longer leading, but still comfortably inside the top ten of a field that is bunched from 1,290 to 1,370 Elo.

Screenshot of the Artificial Analysis Text to Image Leaderboard showing MAI-Image-2.5-Pro at No. 7 with Elo 1,292 behind GPT Image 2 (high), Reve 2.1, Nano Banana 2, GPT Image 1.5 (high), MAI-Image-2.5, and Nano Banana Pro

What MAI-Image-2.5-Pro actually is

MAI-Image-2.5-Pro is Microsoft AI's quality tier of the MAI-Image-2.5 family, announced 23 July 2026 alongside MAI-Voice-2-Flash and placed in public preview on Microsoft Foundry (Azure AI Foundry) and the free MAI Playground. Microsoft describes it as its highest-fidelity image model to date, aimed at hero imagery, detailed editing, and precise in-image text rendering. The family now spans the cost curve: MAI-Image-2.5 for balanced work, MAI-Image-2.5-Flash for high volume, and the Pro tier at the quality end — with a newer sibling, MAI-Image-2.6, already in private preview as of 19 August.

Microsoft's own claims for the Pro tier, all vendor-reported and none independently reproduced yet:

96.8% text-rendering accuracy — flagged as covering Chinese, English, Japanese, Korean and Spanish, though the same documentation caps output at roughly a megapixel, so the "8K" phrasing on the claim is best treated as marketing.

27% better prompt adherence than GPT-Image-1.5 on Microsoft's internal evals.

94% multi-image character consistency across generations.

Up to 84–89% lower GPU cost vs G​PT models in specific Microsoft scenarios (PowerPoint, OneDrive), not a general number.

The documented spec sheet is what grounds those claims: text input up to 32,000 tokens, a 131k context window, 4,096 output tokens, JPEG/PNG image input, and output capped at 1,048,576 pixels (a 1024×1024-equivalent). That last number matters for buyers: this is a high-quality editor, not a print-resolution generator.

Where the quality shows up

The editing strength that the leaderboard is measuring matches what Microsoft emphasizes: precise, surgical edits that keep the rest of the image consistent. Targeted object replacement, layout changes, in-image text updates, and cleanup of artifacts like motion blur — the class of edit where older models regenerate the whole scene and lose the subject. Combine that with the 94% character-consistency claim and you have a model designed for the workflow that matters commercially: take an existing asset, change one thing, ship it.

The #7 text-to-image rank is the honest counterweight. From a blank prompt, MAI-Image-2.5-Pro is good but not the leader — 77 Elo behind GPT Image 2 (high) is a real gap on a blind board. Microsoft is not currently claiming the text-to-image crown; it is claiming the editing crown, and the data supports the narrower claim far better than the broader one.

What it costs

MAI-Image-2.5-Pro is token-metered on Foundry: $5 per million text input tokens, $8 per million image input tokens, $106 per million image output tokens. Microsoft does not publish tokens-per-image, so per-image cost is arithmetic, not a quote; using Artificial Analysis's conversion, a 1024×1024 image lands near $108.50 per 1,000 images. That is roughly 2.3× the sibling MAI-Image-2.5 (~$48 per 1k) and 5.4× MAI-Image-2.5-Flash (~$20 per 1k) — the Pro tier is a deliberate premium.

Preview pricing is not GA pricing; the preview rate card can move before general availability. When it does, a pass-through routing layer is how a price cut reaches your bill the same day: on OrcaRouter, provider list price is passed through at 0% markup, so once MAI-Image-2.5-Pro is reachable through a callable third-party API, whatever Microsoft charges is exactly what you pay — no renegotiation, no second contract.

Scoreboard for MAI-Image-2.5-Pro: image editing rank No. 1 at 1,272 Elo, text-to-image rank No. 7 at 1,292 Elo, price about USD 108.50 per 1k images, 32k text input tokens, 131k context window, ~1-megapixel output cap, preview status on Microsoft Foundry

Why the #1 belongs to Microsoft the company, not just the model

The leaderboard result lands at a moment when Microsoft is visibly replacing third-party image models with its own. MAI-Image-2.5 is already the default engine in Bing Image Creator and runs in PowerPoint, OneDrive, and Dynamics 365 Contact Center; the Pro tier is the same family pointed at higher-fidelity work. Microsoft has said the MAI models are trained on clean, traceable data "without distillation from third-party models" — a direct answer to the OpenAI-dependence question, and the cost claims (up to 84% lower GPU cost in PowerPoint, vendor-reported and scenario-specific) are part of the same story: an in-house model that is cheaper per task than the model it replaces.

None of that would be provable to a skeptical buyer without the independent leaderboard, which is why the #1 is strategically significant rather than just a score. It is the first independent evidence that Microsoft's in-house image line beats the field at the specific task it is built for.

What to watch next

Does the editing lead hold as votes accumulate? 1,272 Elo on ~6,100 samples is a lead, not a lock; Reve 2.1 is 10 points back.

General availability and the final rate card — the preview price is not the GA price.

The resolution ceiling. A 1-megapixel cap keeps Pro out of print work; if Microsoft removes it, the story changes.

MAI-Image-2.6 (private preview, 19 August) — the next family step, ranked No. 2 text-to-image on the other main arena.

Independent evals of the vendor claims — text-rendering accuracy and character consistency are unverified outside Microsoft.

The takeaway from the split board is not "Microsoft now makes the best image model." It is narrower and more useful: Microsoft now makes the best-regarded image editor on an independent leaderboard, at a premium price, in preview, with a hard resolution cap. If your work is changing existing images — hero shots, product visuals, in-image text — that is exactly the model to watch. If your work is blank-canvas generation, the #7 rank says the better pick is still GPT Image 2 — and that is the honest read both numbers give you.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube