Title card for FLUX 3 Image vs Bernini-Diffusers-v2, contrasting 'FLUX 3 Image - live October 1, 2026, $0.048 an image at 1K, on api.bfl.ai' against 'Bernini-Diffusers-v2 - Apache-2.0 weights, self-host only, no hosted API' and noting that OrcaRouter routes neither. The OrcaRouter logo sits bottom-right.
Guides & Insights

FLUX 3 Image vs Bernini-Diffusers-v2: A Priced Endpoint Against a Downloadable Repo

Author

Gideon Frost

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

FLUX 3 Image and Bernini-Diffusers-v2 are both image models you can obtain in full today, and neither of them is a product you can rent. FLUX 3 Image went live on October 1, 2026 at api.bfl.ai/v1/flux-3-image with a published rate card and no published score. Bernini-Diffusers-v2 landed on Hugging Face on August 13, 2026 under Apache-2.0 with published scores and no way to call it except your own GPUs. One comes with an invoice. The other comes with a Python version pin.

That split is the whole comparison, and it is worth being precise about it, because the usual "hosted vs open weights" framing does not fit here. Bernini is not an open-weights model you can drop into an existing inference stack. It is a two-stage system — a Qwen2.5-VL-7B planner in front of a Wan2.2-T2V-A14B diffusion decoder — and its v2 release is a training-schedule fix, not a new capability tier. FLUX 3 Image is a single endpoint with an unusual amount of control surface bolted to it: spatial grounding, a 0–1000 coordinate grid, and box-scoped editing where you name which regions survive. Different shapes of thing.

What each one actually is

• FLUX 3 Image — Black Forest Labs' image tier of its unified FLUX 3 family, released October 1, 2026. One HTTP POST to https://api.bfl.ai/v1/flux-3-image with an x-key header; only prompt is required. The call returns a polling URL, and a finished 4K render can take several minutes. Generated samples are downloadable for one hour.

• Bernini-Diffusers-v2 — ByteDance's planner-plus-renderer stack, pushed to Hugging Face on August 13, 2026 with no accompanying announcement. The repository is tagged image-text-to-video and depends on diffusers. Tasks listed: text-to-image, image-to-image, text-to-video, video-to-video, reference-to-video and reference-to-image. There is no hosted inference and no price list — the getting-started path is hf download and a Gradio app.

• The distribution gap — FLUX 3 Image is a metered service you call. Bernini is a repository you deploy. Before any quality question is worth asking, one of these requires a Hopper-class GPU and the other requires a credit card.

Licensing is where these two diverge hardest

Bernini-Diffusers-v2 is Apache-2.0 at the repository level. For a self-hosted pipeline that matters enormously: no per-image meter, no negotiated commercial agreement, no usage reporting. You pay for electricity and scheduler time, and the marginal cost of the ten-thousandth image is the same as the first.

FLUX 3 Image is not open weights, and Black Forest Labs is explicit that this is temporary. Commercial weights are available today under a negotiated BFL licence, and the public open-weights release is described as coming in the following weeks. That is a promise with a shape but no date, and it is the single most important thing on this page to re-check rather than assume. If you are choosing an image stack for the next two years, the licence question probably outweighs the Elo question — but only if you are prepared to operate a two-stage video-native diffusion stack yourself.

Control surface: bounding boxes versus prompt enhancement

This is the part of the matchup that is genuinely interesting, because the two models solve "make it go exactly there" in completely different ways.

FLUX 3 Image takes a JSON array appended to the prompt. Coordinates run on a 0–1000 grid on both axes, and each box is written as [top, left, bottom, right]. You supply a table of elements, each with an id, a bbox and a desc, plus a global caption that cites those ids by name. In BFL's own example the ids read like animal_1 and silhouette_1; the caption refers to them directly. Above the generation call there is grounding, on by default, which runs a web and image search before generating.

FLUX 3 Image then extends the same coordinate system into editing. Instead of sending a mask, you send a set of box-scoped operations: Keep (source box to the same box, sourced from ref_image_0), Move (source box to a new target box), New (no source, target box generated from a description — add, replace or recolour) and Remove (source box to a null target). Multiple operations go in one request.

The qualification matters. BFL's documentation says pixels outside the edited boxes "usually stay identical." Usually. It is a preservation claim with a hedge built into the wording, which is the honest way to ship one and also the reason you should test on your own frames rather than trusting a demo.

Bernini has no equivalent user-facing coordinate language. Its reference-guided editing improved in v2 because of a training change — the connector module is warmed up for thousands of steps before co-training begins — and ByteDance reports that change showing up as better reference-guided editing and OpenS2V performance. It is a quality fix inside an existing architecture, not a new control interface. If your workflow depends on telling the model which rectangle to leave alone, only one of these two answers the question.

Screenshot of Black Forest Labs' FLUX 3 Image model page, headed 'Maximum control over every pixel', showing the bounding-box and region-editing feature list for the October 1, 2026 image release

What the numbers do and do not support

Bernini-Diffusers-v2's page carries a seven-metric table comparing v2 against v1, and every figure on it is vendor-reported. The direction is mixed, which is worth saying out loud:

• Up — OpenS2V 63.83 (from 62.30), VBench 84.46 (from 84.37), Bernini-rv2v 3.55 (from 3.48).

• Flat — EditVerse 8.02, unchanged, and Bernini-v2v 3.49, unchanged.

• Down — OpenVE 3.96, from 4.03.

The page also states that v2 sits in the first tier among leading closed-source commercial models in an internal blind human pairwise arena. That is unverifiable from outside ByteDance and should be read as a marketing claim, not a measurement. The three real movements are the ones listed above, and they are self-reported against the same lab's own v1.

FLUX 3 Image has no numbers at all yet. Artificial Analysis's text-to-image board and its editing board were both checked on October 1 and 2026-10-02 and carry no FLUX 3 Image row — the BFL entries on those boards stop at the FLUX.2 line, where FLUX.2 [max] sits at 1021 Elo on the generation board and 996 on the editing board. FLUX.2 [dev] anchors both boards at exactly 1000. So the honest statement about FLUX 3 Image quality right now is that nobody outside Black Forest Labs has published a score for it, and the nearest available reference point is the generation before it.

That asymmetry is the practical trap in this comparison. Bernini's numbers are vendor-run but they exist and they are directional. FLUX 3 Image's are absent. Absence is not failure — the model is a day old — but it does mean the only evidence for FLUX 3 Image's editing claim currently comes from the company selling it.

The price side, which is not symmetric at all

FLUX 3 Image is priced per image by resolution tier, and the vendor's rate card lists five of them:

• 768sq — $0.041 list, $0.0205 during the launch window.

• 1K (the default) — $0.048 list, $0.024 on offer.

• 1.5K — $0.07 list, $0.035 on offer.

• 2K — $0.10 list, $0.05 on offer.

• 4K — $0.607 list, $0.3035 on offer.

Screenshot of Black Forest Labs' pricing page captured October 2, 2026, showing the FLUX 3 Image tab and its per-image rates by resolution tier, from 768sq through 4K

Reference images and prompt length are included at no extra charge, and rejected, failed or moderated requests are not billed — which matters more than usual here, because a 4K call is slow and the temptation to abandon a request is real. The launch pricing runs from 15:00 UTC on October 1 to 15:00 UTC on October 8. The list-price jump from 2K to 4K is a factor of 6.07, far steeper than the pixel-count ratio, so the resolution tier you pick is the main cost decision in the whole API.

The tiers track a pixel budget rather than a fixed frame. BFL's example output is 5456 × 3072, roughly 16.8 megapixels, against UHD 2160p at 8.3 megapixels — so "4K" here names a budget, and output dimensions are rounded to multiples of 16 with the actual figure returned as output_mp.

Bernini's price is zero per image and non-zero per hour. Eight GPUs for sequence-parallel inference, Hopper-class silicon recommended, Python 3.11.2 with PyTorch 2.7.1 and CUDA 12.6. The break-even against $0.048 an image at 1K arrives somewhere in the thousands of images per month, depending entirely on what you pay for compute — and it is a real number, not a rhetorical one. If you already run inference infrastructure, Bernini's marginal image is close to free. If you do not, the FLUX 3 Image endpoint is dramatically cheaper than standing up a two-stage diffusion stack.

Routing: neither of these is on our catalogue

Worth stating plainly, because it is the kind of claim that goes stale fast. OrcaRouter does not route FLUX 3 Image, and it does not route Bernini-Diffusers-v2. Calling FLUX 3 Image means calling Black Forest Labs directly at their own endpoint. Calling Bernini means deploying it yourself.

What an OpenAI-compatible router is actually for, in a pipeline like this, is everything on either side of the image call. A captioning or prompt-rewriting pass, a vision model checking the render against the brief, a video model animating the approved still, a small text model classifying the output — those are the steps where being able to swap a provider with one string, on one key, and have requests fail over automatically when one is degraded, removes a real amount of maintenance. OrcaRouter passes provider list price through with no markup, so when a vendor cuts a rate the change is live the same day rather than waiting for a reseller to reprice, and the routing DSL lets you pin a specific model or define fallback order without touching application code. None of that makes FLUX 3 Image routable. It just means the rest of the pipeline does not need to care which model produced the still.

Screenshot of the Artificial Analysis text-to-image leaderboard, showing FLUX.2 [max] at 1021 Elo and FLUX.2 [dev] fixed as the 1000 anchor, with no FLUX 3 Image entry anywhere on the board

Which one to build on

Three questions separate these cleanly, and they do not have a shared answer.

Do you need to control where things go? FLUX 3 Image is the only one of the two with a coordinate language, and its box-scoped edit operations — keep, move, add, remove — are a genuinely different interface from text-instruction editing. If your work is compositing, layout or product imagery where regions must survive, that is decisive and the hedge in "usually stay identical" is something you test rather than assume.

Do you need to know how good it is before you commit? Bernini has numbers, all vendor-reported, and a mixed direction of travel. FLUX 3 Image has none. If you cannot run your own evaluation, you are choosing between a measured model and an unmeasured one, and the measured one is the older architecture.

Can you operate a two-stage diffusion stack? That is the real gate on Bernini. A Qwen2.5-VL-7B planner feeding a Wan2.2-T2V-A14B decoder, on Hopper GPUs, with a prompt-enhancer hook that points at an OpenAI-compatible endpoint. The licence is permissive and the compute bill is not. If you cannot, the question is moot and $0.048 an image is the answer.

The one thing that would change this page fastest is BFL opening the FLUX 3 Image weights, which the company has said is weeks away. That turns the licence asymmetry from decisive to mild and leaves the control-surface difference as the actual comparison. Until then, the accurate summary is that these are two good models with opposite distribution models, and the choice is made on licensing and control long before it is made on quality.