Hero card for the matchup between MAI-Image-2.5-Pro, a premium editing model in Microsoft Foundry preview, and Bernini-Diffusers-v2, ByteDance's Apache-2.0 open-weights pipeline
Guides & Insights

MAI-Image-2.5-Pro vs Bernini-Diffusers-v2: A Measured #1 Against an Unmeasured Open-Weights Bet

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Comparing MAI-Image-2.5-Pro with Bernini-Diffusers-v2 is not really comparing two image models. It is comparing two different answers to the same question — "how do I get a good edit of an existing image?" — where one side has 6,100 blind human votes behind its #1 editing rank and the other has a README. MAI-Image-2.5-Pro, Microsoft's quality-tier image model, went into public preview on Microsoft Foundry on 23 July 2026 and now tops the Artificial Analysis Image Editing Arena. Bernini-Diffusers-v2, ByteDance's second-generation unified generation framework, landed on Hugging Face on 13 August 2026 as Apache-2.0 weights with no announcement and no image-quality benchmark anyone outside ByteDance has run. Every practical difference in this matchup follows from those two facts.

One you call, one you run

MAI-Image-2.5-Pro — a hosted API model in public preview on Azure AI Foundry and the MAI Playground. Proprietary, token-metered, no fine-tuning, no self-hosting. You get an endpoint and pay per token; Microsoft runs the infrastructure.

Bernini-Diffusers-v2 — Apache-2.0 weights you download from Hugging Face and run yourself. The stack is a Qwen2.5-VL-7B semantic planner driving a Wan2.2-based DiT renderer, and the README's hardware recommendation is Hopper-class GPUs (H100/H800/H200) with Python 3.11.2 / PyTorch 2.5.1+cu124. You bring the cluster.

That is the decision rule in one line: if your organization wants to own the stack, Bernini is an option; if your organization wants an invoice and an SLA, only MAI-Image-2.5-Pro is even in the running. They are not substitutes for each other.

The evidence gap is the real headline

The two models sit on opposite sides of the biggest quality problem in image AI: measuring the output. MAI-Image-2.5-Pro's editing skill is independently scored — No. 1 on the Artificial Analysis Image Editing Arena at Elo 1,272 from ~6,100 blind comparisons, ahead of Reve 2.1, GPT Image 2, and its own sibling MAI-Image-2.5. Its text-to-image rank is seventh at 1,292 Elo, which is the honest counterweight: it is a specialist editor, not the blank-canvas leader. Those numbers are auditable by anyone with a browser.

Bernini-Diffusers-v2 has no equivalent public score. The model card reports EditVerse 8.02, OpenVE 3.96, OpenS2V 63.83, and VBench 84.46 — all vendor-reported, and all video-side metrics. There is nothing on the card that an image buyer would recognize as a text-to-image or image-editing leaderboard result. The card does claim Bernini reaches "the first tier" of closed commercial models in ByteDance's internal blind pairwise arena — which is exactly the kind of vendor self-assessment that needs an independent check before it means anything.

MAI-Image-2.5-Pro: editing Elo 1,272 (independent, ~6,100 votes); text-to-image Elo 1,292 (independent, ~11,500 votes)

Bernini-Diffusers-v2: no public image benchmark; video-side metrics only, vendor-reported

Comparison scoreboard for MAI-Image-2.5-Pro and Bernini-Diffusers-v2: editing rank No. 1 at 1,272 Elo versus no public image benchmark; availability Foundry preview versus Apache-2.0 self-host; text rendering 96.8% claimed versus unpublished; price USD 108.50 per 1k versus free weights plus your own compute; editing approach surgical edits versus plan-then-render; scope image-only versus unified image and video

Two different theories of editing

Both models are built around the idea that editing should not mean regenerating the scene. But they get there differently.

MAI-Image-2.5-Pro is Microsoft's "precise, surgical edits with consistency" pitch: change one object, update the in-image text, clean up the motion blur, and keep everything else — subject identity, lighting, texture — stable. That consistency is the mechanism behind its 94% multi-image character-consistency claim (vendor-reported, unverified). The product goal is a model that does the narrow edit and stops.

Bernini-Diffusers-v2 is a plan-then-render pipeline: the Qwen2.5-VL planner decomposes an instruction into semantic steps — change the weather, swap the style, insert an object, keep the subject — and the renderer draws the result. The design is aimed at complex, multi-step instructions, and the same weights cover text-to-image, image-to-image, and video tasks (t2i, i2i, t2v, v2v, rv2v, r2v). That breadth is the upside; the unmeasured cost is that a 7B planner is a real bottleneck for very subtle edits, and nobody outside ByteDance has quantified how it handles them.

Screenshot of the ByteDance/Bernini-Diffusers-v2 model card on Hugging Face showing the README, the Apache-2.0 license, and the benchmark table with video-side metrics

Text rendering, the practical battleground

In-image text is where a modern image editor earns its price, and here the asymmetry is stark. MAI-Image-2.5-Pro's headline capability is precise text rendering — Microsoft claims 96.8% accuracy across English, Chinese, Japanese, Korean, and Spanish (vendor-reported; the same docs cap output near one megapixel, so treat the number as directional). Bernini-Diffusers-v2 has published no text-rendering accuracy at all. For a workflow that produces localized ads, packaging, or anything with a legible word in it, one candidate offers a measurable claim and the other offers nothing.

Cost, and the shape of the bill

MAI-Image-2.5-Pro — $5 per million text-input tokens, $8 per million image-input tokens, $106 per million image-output tokens. Per-image cost is not published; Artificial Analysis's conversion puts a 1024×1024 image near $108.50 per 1,000. Preview pricing, subject to change at GA.

Bernini-Diffusers-v2 — the weights are free, but self-hosting means H100-class hardware (or cloud rental), the ops to keep it running, and your own scaling. The economics invert the API model: expensive for ten images a day, cheap at serious volume.

For most teams the honest framing is: Bernini's total cost of ownership is only favorable if you already run a GPU fleet and the workload is large and steady. Otherwise a metered API at a premium rate is the cheaper failure.

Screenshot of Microsoft's announcement of MAI-Image-2.5-Pro and MAI-Voice-2-Flash, the subject model's vendor page, showing the preview launch and family context

What a router can and cannot do here

Neither model is routable through OrcaRouter today, for opposite reasons: MAI-Image-2.5-Pro sits behind Microsoft Foundry's direct preview, and Bernini-Diffusers-v2 is not an endpoint at all — it is weights. That matters as a principle for this matchup: the router pattern only applies to the half of this comparison that is a callable API. When MAI-Image-2.5-Pro reaches a callable third-party API, the way to adopt it without betting a production path on preview code is to route a slice of traffic to it with a proven model as automatic failover — on OrcaRouter that is one endpoint, provider prices passed through at 0% markup, and a fallback that catches whatever the low-vote leaderboard could not see. The open-weights side has no equivalent: you either run Bernini's stack yourself or you do not use it.

The verdict

MAI-Image-2.5-Pro is a premium, measurable, hosted editor: #1 on an independent editing board, a real text-rendering claim, and a price to match. Bernini-Diffusers-v2 is a free, broad, unproven-in-image open-weights pipeline that also does video — genuinely interesting if you self-host, genuinely unscoreable until someone runs a public image eval. If your workflow needs a callable API and a number you can defend, the choice is MAI-Image-2.5-Pro or one of the other hosted editors it beats. If your workflow needs ownership and you can run the hardware, Bernini-Diffusers-v2 is worth testing — but buy it as a bet on unmeasured potential, not as a competitor to a measured #1.

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube