A hero title card for the comparison 'GPT-Image-2.5 vs Qwen-Image-3.0' with the subtitle 'Two closed flagships, two different price tags', showing a left rounded card labeled 'GPT-Image-2.5 · $8 / $30 per M tokens' with a metering-gauge icon and a right rounded card labeled 'Qwen-Image-3.0 · from 0.18 CNY per image' with a coin icon, and a small 'no weights' badge between the two cards; the OrcaRouter logo is bottom-right.
Guides & Insights

GPT-Image-2.5 vs Qwen-Image-3.0: Two Closed Flagships, Two Different Price Tags

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The clearest way to read GPT-Image-2.5 against Qwen-Image-3.0 is as two flagships that closed their doors in the same summer. On September 8 OpenAI shipped GPT-Image-2.5, the ChatGPT Images 2.5 model served through the API as GPT-Image-2.5 Flare and GPT-Image-2.5 Sunburst, as a hosted-only product priced in tokens — $8 per million image-input tokens and $30 per million image-output tokens. Roughly five weeks earlier, on August 4, Alibaba's Qwen team brought Qwen-Image-3.0 to general availability on Qwen Cloud and Alibaba Cloud Bailian — and this time, in a quiet reversal that matters more than any single feature, it shipped no weights at all. Qwen-Image 1.0 and 2.0 were Apache-2.0 open releases with same-day technical reports; Qwen-Image-3.0 has no downloadable weights, no technical report, no model card, no parameter count, and no benchmarks from its own maker. Both of these are now closed, hosted image flagships from the two labs with the strongest open-weight reputations on either side of the Pacific. The comparison between them is partly about image quality and mostly about what each lab decided to stop giving away.

Neither model is validated by its own maker's documentation, and that is precisely where the two diverge. GPT-Image-2.5 has no independent score yet — it has no Artificial Analysis entry as of September 9, 2026, so its quality claims are OpenAI-reported and unreproduced. Qwen-Image-3.0 ships with nothing from Alibaba, but its Pro edition has nevertheless appeared on the independent boards: on the September Artificial Analysis snapshots, Qwen-Image-3.0-Pro sits mid-table at roughly 1,086 Elo for text-to-image and 1,079 for image editing, listed near $40 per 1,000 images. That is a genuinely independent signal, and it is not a flattering one — it places Alibaba's flagship around ninety Elo behind the leader, GPT Image 2 (high), and below Meta Muse Image and Nano Banana 2. Every price in this piece is a vendor list price; every score is Artificial Analysis's, not Alibaba's.

Why the closed-ness is the story

For Alibaba, closing Qwen-Image is a strategic about-face with a clear logic: the previous open image models were widely downloaded but commercially valueless to Alibaba, because anyone could self-host them. Qwen-Image-3.0 is built to be the opposite — the Qwen team explicitly positions it around the character 实, "practical," meaning production work that runs through a paid hosted API. The same logic already governs the top of the rest of the Qwen family, and it turns every prior "just download Qwen-Image" tutorial into a historical document.

OpenAI never had an open image flagship to close, so its posture is unchanged: GPT-Image-2.5 exists to be called, billed, and metered, with C2PA provenance and watermarking built in. The practical consequence for a developer is identical on both sides of this comparison — there is no self-host option, no fine-tuning of the base, no offline fallback. You are renting the model from its owner on the owner's price sheet, and the only question left is which rental you need.

What each flagship was actually built to draw

Qwen-Image-3.0's engineering story is documents, layout, and language density. Alibaba claims a 4,500-token prompt capacity — roughly four and a half times the prior generation — which is what lets the model compose dense one-pass layouts: full newspaper pages, infographic grids, storyboards, exam papers, and UI mockups. It claims legible text rendering down to about ten pixels including math and LaTeX, native rendering in twelve languages with more than twenty fonts and a hundred artistic styles, and up to 2048×2048 output with as many as six images in a single call. Generation and image-to-image editing live in one model. If the claims hold, it is the strongest hosted model for text-dense, layout-heavy, multilingual image work — and the AA placement of its Pro edition suggests the reality trails the ambition: competent, mid-pack image quality with no sign yet of a text-layout leader.

GPT-Image-2.5's engineering story is the opposite end of the image spectrum: the edit and the subject, not the page. OpenAI's claims center on reference fidelity — keeping a person, pet, or product recognizable across changes of setting and style — multi-turn editing that touches only what was asked, transparent backgrounds, and sharper fine detail, with the Sunburst endpoint dedicated to precision edit work and Flare to fast generation. Its text-rendering claim is narrower than Qwen's (correct handling of Chinese among other scripts) because OpenAI is not chasing document layout; it is chasing the commercial photography and product-design workflow where the single image, edited repeatedly, is the deliverable.

A two-column comparison scoreboard for 'GPT-Image-2.5 vs Qwen-Image-3.0': GPT-Image-2.5 rows read Release GA Sept 8 2026 / Openness closed API / Price $8 in · $30 out per M tokens / Built for editing precision + reference fidelity / Independent signal none yet, no AA entry / Output up to 3840x2160; Qwen-Image-3.0 rows read Release GA Aug 4 2026 / Openness closed, no weights / Price from 0.18 CNY per image, Pro about $40 / 1k on AA / Built for dense layouts, long prompts, small text / Independent signal Pro mid-pack on AA / Output up to 2048x2048, six images per call; footer 'Prices vendor-listed; scores per Artificial Analysis Sept 2026.'; the OrcaRouter logo is bottom-right.

The price comparison neither vendor finishes

Qwen-Image-3.0 lists in per-image prices: domestic consumer pricing starts around ¥0.18 per image for the Standard edition, with Pro higher — roughly ¥0.25 at 1K resolution and up to about ¥0.5 at 2K — and international Standard pricing in the $0.03–$0.04 per-image range. Independent boards record the Pro edition near $40 per 1,000 images. GPT-Image-2.5 lists in tokens and publishes no per-image equivalent; its rate card matches GPT-Image-2, which Artificial Analysis anchors near $211 per 1,000 images at high quality. Even granting that the OpenAI figure is for the predecessor, the list-price gap is on the order of five to eight times, and it buys two completely different quality promises: Qwen's promise is measured in the volume of correctly rendered small text on a dense page, OpenAI's in the fidelity of an edited subject.

The comparison is also one neither vendor lets you complete cleanly. Qwen's Standard pricing is real but the Standard edition's quality is a black box — Alibaba publishes nothing, and only the Pro edition's mid-pack AA scores exist to anchor the family. OpenAI's quality claims are specific and internally evaluated, but its per-image cost is unknown and its output is independently unmeasured. You can price both families and you can only partially validate either one.

Dimension by dimension

Release — GPT-Image-2.5: GA September 8, 2026. Qwen-Image-3.0: GA August 4, 2026 (announced July 21; Pro edition independently listed since July).

Openness — Both closed, hosted-only. Qwen-Image-3.0's closure reverses Apache-2.0 open releases of Qwen-Image 1.0 and 2.0; GPT-Image-2.5 was never open.

Price — GPT-Image-2.5: $8 image-in / $30 image-out per million tokens, per-image unpublished. Qwen-Image-3.0: from ~¥0.18 per image (Standard, ~$0.03 international); Pro ~¥0.25–0.5, ~$40/1k on AA.

Strength — GPT-Image-2.5: editing precision, reference fidelity, multi-turn consistency. Qwen-Image-3.0: dense document/layout generation, 4,500-token prompts, ~10px legible text, 12 languages.

Independent signal — GPT-Image-2.5: none yet. Qwen-Image-3.0: Pro mid-pack on AA — ~1,086 T2I, ~1,079 editing, roughly 90 Elo behind GPT Image 2 (high).

Output — GPT-Image-2.5: up to 3840×2160, transparent backgrounds. Qwen-Image-3.0: up to 2048×2048, up to six images per call.

Which rental fits your workload

If your work is documents, education, localization, or any high-volume multilingual layout — course materials with equations, newspaper and magazine pages, UI mockups, infographics that must hold legible text in a dozen languages — Qwen-Image-3.0 is the flagship built for that job, and Standard's per-image pricing makes volume experiments cheap. The condition is trust: the Standard edition is undocumented, and the independent scores that exist are for the pricier Pro edition, which lands mid-pack — fine for layout work, not a quality leader. Budget for visual QA at scale.

If your work is commercial imagery — product shots, campaign creative, reference-consistent edits where the subject must survive unchanged across a long edit chain — GPT-Image-2.5 is the better-shaped bet, with the same caveat in a milder form: at least OpenAI publishes a rate card and specific claims, and its predecessor's measured #1 standing on the AA text-to-image board gives you a rational prior. Neither model is on OrcaRouter's catalog as of September 9, 2026 (the prior OpenAI generation, GPT-Image-2, is, at provider list price with no markup), so an A/B trial means holding keys for both Qwen Cloud and the OpenAI API. The pass-through idea still applies once these do route: an aggregator that bills provider list price with no markup is the cleanest way to run cross-vendor trials like this without building per-provider billing.

Screenshot of the Artificial Analysis text-to-image leaderboard captured September 9 2026: GPT Image 2 (high) is ranked first at Elo 1,178; Qwen-Image-3.0-Pro appears around twelfth at about Elo 1,086, roughly ninety Elo behind the leader and below Meta Muse Image and Nano Banana 2; GPT-Image-2.5 has no entry.

The open question both models leave is whether "closed flagship" is a durable strategy for image generation or a temporary one. Qwen-Image-3.0's closure is the more reversible — Alibaba has precedent for open and a community that expects it — and the louder signal when it ships weights again will be that the hosted-only experiment failed to earn enough. GPT-Image-2.5's bet is that OpenAI's editing lead, once measured, will justify token-metered prices against per-image competitors. Watch for two events: Qwen-Image-3.0 (non-Pro) appearing on an independent board at all, and GPT-Image-2.5 doing the same. Until then, this is a matchup of two price sheets attached to quality promises only one of which — the Qwen Pro edition — has so far been checked by anyone independent, and that check came back mid-pack.

Screenshot of the Artificial Analysis image-editing leaderboard captured September 9 2026: MAI-Image-2.6 is ranked first at Elo 1,122, GPT Image 2 (high) second at 1,117, and Qwen-Image-3.0-Pro appears lower mid-table at about Elo 1,079, with its API price listed near $40 per 1,000 images.