
FLUX 3 Image vs FLUX 3: Two Names, Two Products, Nine Weeks Apart
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 223 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 125 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 103 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 212 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
Since October 1, 2026, "FLUX 3" has meant two different things, and the second meaning is the one with a price list. FLUX 3 Image is the image tier that went live that day at api.bfl.ai/v1/flux-3-image, metered per image from $0.041. FLUX 3 is the larger thing it belongs to — the unified multimodal model Black Forest Labs announced on July 23, 2026 and has since shipped in pieces, of which video and a robotics action head are the parts already running in production. This is not a comparison between a flagship and its smaller sibling. It is a comparison between a whole family and one of its three working parts.
That distinction is easy to lose, and BFL's own naming does not help: the family page, the video documentation, the license terms and the API reference all say FLUX 3 in places where they mean something narrower. If you have read a FLUX 3 spec sheet and come away unsure whether it describes an image model, a video model or a research programme, that is not you being careless. Both names are in active use for overlapping subsets of the same release.
The two dates that matter
• July 23, 2026 — FLUX 3 announced under an Early Access banner. One Self-Flow architecture spanning image, video and audio generation plus action prediction. The compute disclosure is the interesting part: more than 95% of training compute went to video, and fewer than 0.5% of tokens were allocated to audio.
• October 1, 2026 — the image tier opens as a metered endpoint. Which leaves a gap of roughly ten weeks between the announcement and the arrival of the modality most people assumed "FLUX 3" referred to.
Reading that compute split tells you what Black Forest Labs believed the hard problem was. A model family that spends almost all of its training budget on video is not optimising for still images and treating motion as an extension. It is the reverse, and the image tier inherited whatever the video-first training produced.

What each tier is, and what it costs
Three of the four modalities announced in July have shipped. The image tier is the newest, and it is the only one that behaves like a conventional paid API.
• FLUX 3 Image — one POST, an x-key header, and a polling URL back. Priced per image by resolution tier: $0.041 at 768sq, $0.048 at 1K (the default), $0.07 at 1.5K, $0.10 at 2K and $0.607 at 4K. A launch discount of 50% runs until 15:00 UTC on October 8. References and prompt length are included; failed and moderated requests are not billed. It accepts up to ten reference images between 256×256 pixels and 16 megapixels each.
• FLUX 3 Video — generally available and metered by the second rather than the image. Text-to-video and image-to-video run $0.06/s in Draft HD, $0.17/s HD, $0.29/s FHD, $0.40/s QHD and $0.80/s UHD, with video-to-video higher across the board. Clips run to 20 seconds with native synchronised audio, and the modes include keyframe-to-video and continuation. Draft mode is HD-only. QHD and UHD output arrived on September 10, 2026.
• FLUX 3 Action — a 7B world-action model, and the only part of FLUX 3 whose parameters you can download. Three checkpoints appeared on Hugging Face on September 22, 2026: a base checkpoint, a SO-101 arm variant and a DROID variant. Access is partner-gated rather than open, with robotics manipulation as the demonstrated use.
• What has not shipped — a FLUX 3 open-weight backbone. BFL's self-hosting tiers still cover only FLUX.2 [klein] and FLUX.2 [dev]. For the image tier specifically, commercial weights exist under a negotiated licence and public open weights are described as weeks away.
The cost comparison nobody makes

Per-image and per-second pricing look incomparable, which is exactly why the arithmetic is worth doing once.
Five seconds of standard HD text-to-video is $0.85. Five seconds at UHD is $4.00. A single 4K still from the image tier is $0.607 list, or $0.3035 inside the launch window — and while the image tier's "4K" names a pixel budget rather than a fixed frame (BFL's example output is 5456 × 3072, about 16.8 megapixels, against UHD 2160p's 8.3), it is the same order of resolution the video tier charges by the second to move.
The consequence is that generating one UHD keyframe costs a fraction of generating one second of UHD video, and the keyframe is the thing you would generate repeatedly while iterating. If your pipeline is "get the frame right, then animate it," the image tier is the cheap half by more than an order of magnitude, and it is the half that has been available for the shortest time.
There is a caveat on the image side that the video side does not have. A 4K image call can take several minutes, and the returned sample is downloadable for one hour. Slow and one-hour-retrievable is a different operational shape from a video job, and it is the kind of detail that decides whether a batch pipeline needs a queue.
Bounding boxes exist in one of these two names
The control surfaces are not shared across tiers, and this is the sharpest practical difference between the name and the family.
The image tier takes a JSON array appended to the prompt, with coordinates on a 0–1000 grid on both axes and each box written [top, left, bottom, right]. You pass a table of elements carrying an id, a bbox and a desc, and a global caption that cites those ids — BFL's examples use names like animal_1. The same coordinate system drives editing: keep a region, move a region, generate a new one into a target box, or remove a source box. Multiple operations go in a single request, with ref_image_0 naming the first entry in the reference list. Pixels outside the edited boxes, the documentation says, "usually stay identical" — a hedge worth reading literally.
Grounding is on by default in the image tier: a web and image search runs before generation. Output dimensions are rounded to multiples of 16 and the realised figure comes back as output_mp.
None of this appears in the video tier's parameters. FLUX 3 Video has its own control story — keyframes, continuation, first and last frame conditioning — but no coordinate language. The parts share an architecture, not an interface.
Reading the vendor's own comparisons carefully
Black Forest Labs published preference results for FLUX 3 Video, and they are worth quoting with their provenance attached. The runs are vendor-executed, preliminary, and scored as win rates on 10-second 720p output:
• Against Luma Ray 3.2 — 93% preference for FLUX 3 Video.
• Against Runway Gen-4.5 — 77%.
• Against Grok Imagine Video — 69%.
• Against Kling v3 Pro — 60%.
• Against Happy Horse v1 and v1.1 — 59% and 57%.
• Against Seedance 2.0 and Gemini Omni Flash — roughly 52% each, which is a statistical tie in substance even if the number is above half.
The spread from 93% down to 52% across that list is more informative than any single row. Nobody outside BFL has run these comparisons, and the sub-55% results at the bottom are the ones that tell you where the model is actually contested.
There are no equivalent numbers for FLUX 3 Image at all. Artificial Analysis's text-to-image and editing boards were checked on October 1 and again on October 2, 2026, and neither carries a FLUX 3 Image entry. The FLUX.2 line is still where BFL's published scores stop: FLUX.2 [max] at 1021 Elo on the generation board and 996 on the editing board, with FLUX.2 [dev] fixed as the 1000 anchor on both. The image tier is one day old and unmeasured.
![Screenshot of the Artificial Analysis text-to-image leaderboard captured October 2, 2026, listing FLUX.2 [max], FLUX.2 [flex] and FLUX.2 [dev] as the 1000-point anchor, with no FLUX 3 Image row present](https://cms.orcarouter.ai/api/media/file/4-1496.png)
Choosing between the names
The choice is less about which model is better than about which of the three shipped tiers your problem lives in.
If you need stills with spatial control, FLUX 3 Image is the only tier with a coordinate language, and at $0.048 an image at 1K it is cheap enough to evaluate properly rather than argue about. Start there, test the "usually stay identical" claim on your own frames, and watch the 2K-to-4K list jump of 6.07× when you size your batches.
If you need motion, FLUX 3 Video is a mature metered product with resolution tiers up to UHD and 20-second clips carrying synchronised audio. Compare it against the models it is built to displace rather than against the image tier — they are not alternatives to each other.
If you need weights, the answer today is FLUX 3 Action under a partner agreement, and otherwise FLUX.2. Neither the video nor the image tier is downloadable, and the image tier's open release is a stated intention without a date.
Neither FLUX 3 nor FLUX 3 Image is on OrcaRouter's catalogue, so calling either one means going to Black Forest Labs directly. What a vendor-neutral router earns in a pipeline built around this family is the layer around the generation call: the prompt-rewriting pass, the vision check that compares a render against the brief, the classifier that sorts outputs into accept-and-reject buckets. Those steps change far more often than the image model does, and swapping a provider behind one OpenAI-compatible endpoint — with provider list prices passed through unmarked-up, automatic failover when a backend degrades, and a routing DSL when you want to pin or fall back deliberately — is the difference between a pipeline that tracks the market and one that has to be redeployed every time a model gets deprecated.
The honest close is a naming convention. When someone says FLUX 3 and hands you a rate card, they mean the image endpoint. When they say it and hand you a spec sheet about audio token allocation, they mean the July announcement. The gap between those two documents is nine weeks and three shipped tiers, and it is still widening.
