
Hy Image3.5 vs LLaDA-Image: Rent the Endpoint or Own the Weights
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens
- deepseekNEWDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
- qwenQwen: Qwen3.8 Max2026-08-0345Intelligence76Coding
- deepseekDeepSeek: DeepSeek V4 Flash 07312026-07-3134Intelligence69Coding
- minimaxMiniMax: MiniMax-H32026-07-31minimax/minimax-h3
- qwenQwen: Qwen3.7 Flash2026-07-27$0.03 / $0.13 per 1M tokens
- orcaOrcaDub: OrcaDub 1.02026-07-27orca/dub
There are two ways to buy a 2K image model in September 2026 and they are not two products so much as two business models. Hy Image3.5 preview, live since 22 September 2026, is a metered endpoint: $0.024 an image on Tencent's international rate card, ¥0.15 domestically, billed on delivered output, and you never touch a GPU. LLaDA-Image, from the inclusionAI open-model group and released on 4 September 2026, is a downloadable checkpoint: a 7B-class unified generation and editing family under an Apache-2.0 tag, sitting on Hugging Face with no hosted endpoint attached to it. A hundred thousand 2K images a month costs about $2,400 the first way. The second way costs nothing per image and an amount of money you have to work out yourself, because the weights are free and the hardware is not.
That is the whole comparison, and almost every mistake made about it comes from comparing the two as if the price column were the only difference. It is not. One of these can be deprecated under you. One of them can be fine-tuned. One of them comes with a vendor's own benchmark and no independent score; the other comes with a technical report, a public repository, and about a thousand downloads.
Two claims attached to LLaDA-Image that do not agree with themselves
Before the comparison is worth anything, two details on the LLaDA-Image card need flagging rather than resolving, because both are genuine conflicts in the public record and quietly picking a side would be a worse error than saying so.
The first is the parameter count. The model card's own prose describes "a competitive 6B-parameter open-source unified image generation and editing model family." The sidebar on that same page reports the Safetensors model size as 7B params at BF16. Both numbers are on the same page, published by the same organisation, and nothing in the card reconciles them. Treat the class as "roughly seven billion parameters, with a six-billion figure in the description" and do not quote either as settled. Anyone telling you the exact count is guessing from the same card you can read.
The second is the licence. The card carries License: apache-2.0 today. Our own earlier coverage of this family said none of the released cards carried a licence tag and that there was no LICENSE file in the repository. Both statements can be true at different times — a tag can be added — but we are not going to pretend they were always the same, and a licence that appeared after release is a materially different fact from a licence that shipped with the weights. If the licence is load-bearing for your use case, check the repository yourself at the moment you download, and check whether the training code, which the card describes as coming soon, arrives under the same terms.

What the two architectures are actually selling
LLaDA-Image is not a diffusion model in the usual sense and that is the interesting part of it. It pairs a diffusion transformer with a frozen vision-language model — LLaDA2.0-mini — and an RQA connector between them, and it is published as a family rather than a single checkpoint: Base at 50 steps for quality, Turbo at two to four steps for throughput, each in BF16 and FP8. That is four checkpoints, and the step counts are the reason a self-hosted deployment can be made to work on a budget at all: a 50-step 7B model is a different hardware proposition from a 2-to-4-step one. The technical report is public and the training recipe is described as fully open, which is a stronger transparency claim than most open-weight releases make.
Hy Image3.5 preview is the opposite shape by design. It is one endpoint with one behaviour: text-to-image and image-to-image on the same call, multi-turn editing that carries earlier turns as context, up to five reference images per request, a 2K output ceiling, and the identifier hy-image-v3.5-preview. There is nothing to choose, nothing to tune, and nothing to download. The capability set is broader out of the box — multi-turn conversational editing with retained context is a real feature that the open family does not advertise — and you have no control over any of it.
• Cost basis — Hy Image3.5 preview is $0.024 per 2K image metered on delivered output; LLaDA-Image is free weights plus whatever your GPU costs to run.
• What you get — Hy Image3.5 preview is a hosted API with multi-turn editing and five reference images; LLaDA-Image is four checkpoints (Base and Turbo, BF16 and FP8) and a training recipe.
• Steps to an image — Hy Image3.5 preview is a single call with no step parameter exposed; LLaDA-Image is 50 steps on Base or 2 to 4 on Turbo, which you set yourself.
• Weights — Hy Image3.5 preview publishes none; LLaDA-Image publishes the full family.
• Licence — Hy Image3.5 preview is a commercial service governed by Tencent's terms; LLaDA-Image carries an apache-2.0 tag that our own earlier reporting did not see.
• Evidence — Hy Image3.5 preview has a vendor GSD blind test and no independent Elo; LLaDA-Image has an author-reported Qwen-Image-Bench figure of 53.53 in English and 53.38 in Chinese, and also no independent Elo.
The things only one of these can do
Start with what the weights buy you, because it is more than a licensing preference. A checkpoint is an asset: it cannot be deprecated, repriced, rate-limited or shut off, and a pipeline pinned to a specific file will still run in three years. It can be fine-tuned on your own material, which for a team with a house style or a specific character set is the difference between prompting someone else's model and training your own. It runs offline, which matters if the material is client work under a confidentiality term. And its throughput is a function of the hardware you bought rather than a queue you joined.
Against that, an endpoint buys you an absence of work. Nobody on your team learns a serving stack, nobody sizes a GPU, nobody discovers in production that the FP8 checkpoint needs a different runtime than the BF16 one, and nobody is on the hook when the job fails at 3am. The Tencent model also carries a capability the open family does not claim: multi-turn editing that retains the conversation, which is precisely the mechanism behind holding a character across a long series of edits. LLaDA-Image's editing support is real — it is a unified generation-and-editing family — but conversational state across turns is not something the card advertises.
The honest framing is that these are not competing on quality at all. They are competing on who carries the operational risk, and the two answers are "Tencent" and "you".
What has actually been measured, on either side
Neither model has an independent placement. Artificial Analysis's text-to-image board, read on 23 September 2026, carries no Hy Image3.5 preview row and no LLaDA-Image row. Where the board does have Tencent is the previous generation: HunyuanImage 3.0 Instruct at 964 Elo and HunyuanImage 3.0 at 945, ranked 65th and 70th of 156 models — which is the only third-party measurement of anything in this family, and it is a generation old.
What each side offers instead is its own work. Tencent's is the GSD-style blind preference test run with several hundred of the company's own professional designers, reporting roughly thirty percent over Hy Image3.0, parity with Seedream 5.0 Pro, and a slight edge over Nano-Banana Pro and Qwen-Image-3.0-Pro. Vendor-run, vendor-populated, prompt set unpublished — a real method with a real conflict of interest, and both halves of that sentence belong in the same breath.
LLaDA-Image's is a Qwen-Image-Bench score of 53.53 in English and 53.38 in Chinese, which is an author-reported number on a benchmark the authors chose. It is not more independent than Tencent's test; it is differently unaudited. The card also lists five community Spaces and a monthly download count around a thousand, which tells you the open family has real but modest traction — it is not a model the field has already adopted, and anyone who tells you otherwise has not looked at the number.

Where the money actually goes, at volume
Run the rent side honestly, because it is the only side with a published rate. A hundred thousand 2K images a month at $0.024 is $2,400. Half a million a month is $12,000, or about $144,000 a year, and at that scale a self-hosted 7B-class image model stops being a hobbyist option and starts being an obvious line item. Below roughly ten thousand images a month the arithmetic does not work that way at all: $240 against the cost of a GPU, the power, the storage, the serving stack and somebody's attention is not a close call, and the metered endpoint wins on every axis including the one you are not counting.
The crossover is not a number anyone can hand you, because the self-hosted side depends on hardware you already own, utilisation you actually achieve, and whether the GPU is doing something else when it is not generating images. What can be said precisely is the shape of the curve: the metered model is linear in volume and the self-hosted model is a step function, so the question is not which is cheaper but where your volume sits relative to the step.
One endpoint, whichever way you go
OrcaRouter routes neither of these. We host no Tencent image model — the only Tencent entries in the catalogue are the text-only Hy3 line — and nothing from inclusionAI at all, so Hy Image3.5 preview means Tencent's own cloud and the first-party Tencent products, and LLaDA-Image means the published weights and whatever runtime you point them at. Saying that plainly is more useful than a model page that leaves it ambiguous.
What a routing layer is good for in this decision is the third option it makes cheap. If you are weighing $2,400 a month against a GPU purchase, the useful intermediate is knowing what the same workload costs across models you can reach today without a contract — openai/gpt-image-2, google/gemini-3.1-flash-image-preview, the Imagen 4 family and grok/grok-imagine-image all sit behind one OpenAI-compatible endpoint at provider list price with 0% markup, so a vendor price change is live on our side the same day, and automatic failover means one provider's bad afternoon is not your outage. That is not a substitute for either model in this comparison. It is what stops the comparison from being a two-way choice made in the dark.

The one question that decides it
Ask whether your workload has a shape that the weights can capture and the endpoint cannot. If you are generating against a house style, a fixed character set, or a client's own material, and you have the volume and the hardware to make a 7B-class model earn its keep, then LLaDA-Image is the more durable asset and the Apache-2.0 tag — check it at download time, given the history above — is what makes it usable. You are buying optionality and paying for it in operations.
If your workload is bursty, if nobody on the team wants to own a serving stack, if multi-turn editing with retained context is the feature you actually need, or if your volume is under ten thousand images a month, then $0.024 a delivered 2K image is a good price for someone else's problem and the preview label is the main thing to keep in view. The one thing worth doing before 7 October 2026 — while Miora, WorkRally and OnSolo are running the free window — is taking your own prompt set through the tencent model and writing down the accept rate, because that number, not the thirty percent, is the one that decides whether the metered side is worth its meter.
