Hero title card for GPT-Image-2 vs Flux 3, headline 'GPT-Image-2 vs Flux 3' with subtitle 'One live model, one roadmap', two cards showing GPT-Image-2 with a badge 'Public API since May 2026' and Flux 3 with a badge 'Image tier: coming soon', OrcaRouter logo in the bottom-right corner.
Guides & Insights

GPT-Image-2 vs Flux 3: Comparing a Live Model to a Roadmap

Author

Alistair Wren

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

This is not a normal model-versus-model post, because GPT-Image-2 and Flux 3 are not in the same state of existence. GPT-Image-2 shipped on April 21, 2026, has been callable through Ope​nAI's API since early May, and gained a new capability this week — transparent-background image generation in the API, announced August 20 and available now. Flux 3 is Black Forest Labs' multimodal foundation model, announced July 23, 2026, and the image tier still has no public endpoint, no model ID, and no price card as of this writing. One of these you can call at 3 a.m. with a valid key; the other you can sign up to be waitlisted for. The useful comparison, then, is not "which is better" — it is what the shipped model actually does today, what the roadmap claims for the unshipped one, and how a builder should treat a promised model that keeps slipping while a live one keeps improving.

One of these has an endpoint

Start with availability, because it decides everything else. Black Forest Labs unveiled Flux 3 on July 23 as a single model trained jointly on image, video, audio, and robot-action prediction, built on a flow-matching approach the lab calls Self-Flow. Over 95% of that training compute reportedly went to video prediction, which is visible in what actually shipped: only Flux 3 Video reached general availability, on August 4, with a priced API. The image tier's model page still carries a "coming soon" label with an application form, not a get-started button — no endpoint, no documented parameters, no per-image cost, and no benchmark anyone can run.

GPT-Image-2, by contrast, has been in production for four months. Ope​nAI opened API access in early May, and the model's identifier, gpt-image-2, is stable enough that a pinned snapshot, gpt-image-2-2026-04-21, coexists with it. It is the current number-one text-to-image model on both major independent boards — 1,381 Elo on Arena's text-to-image leaderboard (the medium build) and 1,369 Elo for the high build on Artificial Analysis — and it renders multilingual text including CJK, grounds generation in web search, and produces up to 4K output. Four months of production traffic is the difference between a claim and a track record.

The recent change on the GPT-Image-2 side

The event that makes this matchup timely has nothing to do with Flux 3. On August 20, Ope​nAI opened a transparent-background preview for GPT-Image-2 in the Images API: set background to "transparent" and the endpoint returns a PNG or WebP with a real alpha channel, cut out and ready to composite. That is a feature aimed squarely at the production work Flux 3 image's roadmap is also selling — product shots, icons, marketing assets — and it landed as a parameter on an already-live endpoint rather than a new model. So while the world waits on a Flux 3 image endpoint that still says "coming soon," the incumbent just closed one of the gaps that kept it out of compositing pipelines.

The capability is API-only — there is no ChatGPT toggle — and it is a parameter, not a new product. Both Images API endpoints, generation and edit, accept a background field with values "transparent", "opaque", and "auto" (the default, where the model decides). The alpha channel survives only in PNG or WebP output, so the request pins the output format explicitly. Ope​nAI's own API reference still marks transparent support for gpt-image-2 and its gpt-image-2-2026-04-21 snapshot as "in preview" — that is the vendor's label, not a graduation announcement — and its guidance is to keep the prompt free of any described backdrop so the model actually honors the transparent request. No new endpoint, no new model ID, no separate price card: the same gpt-image-2 call that returned an opaque PNG now returns a cutout.

Comparison scoreboard titled 'GPT-Image-2 vs Flux 3 - the scoreboard'. GPT-Image-2 column: Released Apr 21, 2026; Availability Public API since May; Arena T2I Elo 1,381 (#1); Text rendering multilingual incl. CJK; Transparent bg preview Aug 20; Max resolution 4K. Flux 3 column: Released announced Jul 23, 2026; Availability image tier coming soon; Arena T2I Elo none yet; Text rendering promised; Transparent bg not stated; Max resolution not specified. Footer: GPT-Image-2 figures per Arena and OpenAI list pricing; Flux 3 image claims are a vendor roadmap. OrcaRouter logo bottom-right.

What Flux 3 image promises, and what is actually shipping

The Flux 3 image tier's spec sheet is a vendor roadmap, so treat it as one. Black Forest Labs has promised image synthesis and editing across styles and resolutions, improved complex-prompt handling, high-accuracy multilingual text rendering, zero-shot style transfer from reference images, character and brand consistency without fine-tuning, and non-destructive editing that preserves texture and lighting. Those are the right things to promise — they are the exact features the current market leaders compete on — but none of them is testable today because none of them has an endpoint.

What is actually callable from BFL is the video tier and the older FLUX image line. Flux 3 Video sells for $0.17 per second at HD and $0.29 per second at FHD for full renders, with a $0.06-per-second draft mode. Its own reported numbers are a 1,135 Elo for text-to-video and 1,051 for image-to-video, and preference wins over Luma Ray 3.2, Runway Gen-4.5, Gr​ok Imagine Video, and Kling v3 Pro — all vendor-reported and unreproduced, and none of it transfers to the image tier anyway. For still images, the current BFL offering remains the FLUX.2 line, which sits at 1,226 Elo on the same Artificial Analysis board where GPT-Image-2 leads at 1,369. That gap is the practical context for "Flux 3 image is coming": today's BFL image API is roughly 140 Elo behind Ope​nAI's, and the promised model is what BFL needs to close it.

Screenshot of the Artificial Analysis Text-to-Image leaderboard (captured August 20, 2026) showing GPT Image 2 (high) in first place at 1,369 Elo and Black Forest Labs' current image entry FLUX.2 (max) at 1,226 Elo, with no Flux 3 image entry present, in English.

Pricing: only one real price card exists

There is no Flux 3 image price, because there is no Flux 3 image product. The only published numbers are GPT-Image-2's token rates: $5 per million text input tokens, $8 per million image input tokens, and $30 per million image output tokens, with the Batch API running at roughly half for up-to-24-hour turnaround. Ope​nAI does not publish tokens-per-image, so per-image figures are conversions, not quotes: roughly $0.006 for a low-quality 1024×1024, around $0.05 for a medium one, around $0.21 at high quality, and up to about $0.41 for a 4K high render. For a comparison, that is the one side of the ledger that is real; the Flux 3 image column is an empty cell.

Screenshot of the OrcaRouter model page for openai/gpt-image-2 (captured August 20, 2026), showing the model description, input rate of $8.00 and output rate of $30.00 per million tokens, median latency, and community usage, in English.

What a builder should do with a promised model

The sane posture when a competitor's model keeps slipping is to build on what exists and make the future a configuration change rather than a re-architecture. GPT-Image-2 is routable today through OrcaRouter — one API key, Ope​nAI's list price passed through at 0% markup, automatic failover across providers — so a pipeline that generates assets with it can later point the same endpoint at Flux 3 image's API the day an endpoint actually exists, and route a slice of traffic to it with the routing DSL without touching calling code. You get the shipped model now, and when the promised one lands you evaluate it against a working baseline instead of against a press release. The wait costs you nothing; building on nothing while you wait costs you everything.

The verdict is lopsided by design. If you need image generation this quarter, GPT-Image-2 is the pick and there is no second question — it is the independent number one, it just added transparent-background output in the API, and it has a public API and a stable price. Flux 3 image is a reason to keep an eye on Black Forest Labs, not a reason to hold a project. Track the early-access opening and the first independent benchmarks; the moment an endpoint exists, the comparison in this post becomes a test you can actually run.