Hero title card reading 'Hy Image3.5 vs LLaDA-Image - the endpoint against the checkpoint', subtitled 'Two ways to buy a 2K image model in September 2026, and who carries the operational risk'. The left column, 'Hy Image3.5 preview - rent', lists: access via a metered endpoint with no weights; $0.024 per 2K image billed on delivered output; a 2K ceiling and up to 5 reference images; multi-turn editing with retained context, yes; fine-tunable, no; can be deprecated under you, yes. The right column, 'LLaDA-Image - own', lists: access via downloadable checkpoints with no hosted endpoint; free weights plus your own GPU; a family of Base 50-step and Turbo 2-4 step in BF16 and FP8; multi-turn editing with retained context not advertised; fine-tunable, yes; can be deprecated under you, no. A footer reads: 'Tencent figures vendor-reported. LLaDA-Image is in inclusionAI's open family; its card carries an apache-2.0 tag and reports 6B in prose against 7B in the sidebar. OrcaRouter routes neither.' The OrcaRouter logo is composited in the bottom-right corner.
Guides & Insights

Hy Image3.5 vs LLaDA-Image: Rent the Endpoint or Own the Weights

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

There are two ways to buy a 2K image model in September 2026 and they are not two products so much as two business models. Hy Image3.5 preview, live since 22 September 2026, is a metered endpoint: $0.024 an image on Tencent's international rate card, ¥0.15 domestically, billed on delivered output, and you never touch a GPU. LLaDA-Image, from the inclusionAI open-model group and released on 4 September 2026, is a downloadable checkpoint: a 7B-class unified generation and editing family under an Apache-2.0 tag, sitting on Hugging Face with no hosted endpoint attached to it. A hundred thousand 2K images a month costs about $2,400 the first way. The second way costs nothing per image and an amount of money you have to work out yourself, because the weights are free and the hardware is not.

That is the whole comparison, and almost every mistake made about it comes from comparing the two as if the price column were the only difference. It is not. One of these can be deprecated under you. One of them can be fine-tuned. One of them comes with a vendor's own benchmark and no independent score; the other comes with a technical report, a public repository, and about a thousand downloads.

Two claims attached to LLaDA-Image that do not agree with themselves

Before the comparison is worth anything, two details on the LLaDA-Image card need flagging rather than resolving, because both are genuine conflicts in the public record and quietly picking a side would be a worse error than saying so.

The first is the parameter count. The model card's own prose describes "a competitive 6B-parameter open-source unified image generation and editing model family." The sidebar on that same page reports the Safetensors model size as 7B params at BF16. Both numbers are on the same page, published by the same organisation, and nothing in the card reconciles them. Treat the class as "roughly seven billion parameters, with a six-billion figure in the description" and do not quote either as settled. Anyone telling you the exact count is guessing from the same card you can read.

The second is the licence. The card carries License: apache-2.0 today. Our own earlier coverage of this family said none of the released cards carried a licence tag and that there was no LICENSE file in the repository. Both statements can be true at different times — a tag can be added — but we are not going to pretend they were always the same, and a licence that appeared after release is a materially different fact from a licence that shipped with the weights. If the licence is load-bearing for your use case, check the repository yourself at the moment you download, and check whether the training code, which the card describes as coming soon, arrives under the same terms.

Screenshot of the Hugging Face model card for inclusionAI/LLaDA-Image, showing License: apache-2.0, tags including Text-to-Image, Diffusers, Safetensors, LLaDAImagePipeline, image-generation, image-editing and arxiv:2609.03796, model-scope badges for Base, Base-FP8, Turbo and Turbo-FP8, the description text 'LLaDA-Image is a competitive 6B-parameter open-source unified image generation and editing model family', and a right rail listing 1,021 downloads last month, Safetensors with model size 7B params at BF16 tensor type, an inference-providers panel reading 'This model isn't deployed by any Inference Provider', a model tree with 1 finetune and 3 quantizations, 5 Spaces, and a collection updated 19 days ago with the paper published 20 days ago.

What the two architectures are actually selling

LLaDA-Image is not a diffusion model in the usual sense and that is the interesting part of it. It pairs a diffusion transformer with a frozen vision-language model — LLaDA2.0-mini — and an RQA connector between them, and it is published as a family rather than a single checkpoint: Base at 50 steps for quality, Turbo at two to four steps for throughput, each in BF16 and FP8. That is four checkpoints, and the step counts are the reason a self-hosted deployment can be made to work on a budget at all: a 50-step 7B model is a different hardware proposition from a 2-to-4-step one. The technical report is public and the training recipe is described as fully open, which is a stronger transparency claim than most open-weight releases make.

Hy Image3.5 preview is the opposite shape by design. It is one endpoint with one behaviour: text-to-image and image-to-image on the same call, multi-turn editing that carries earlier turns as context, up to five reference images per request, a 2K output ceiling, and the identifier hy-image-v3.5-preview. There is nothing to choose, nothing to tune, and nothing to download. The capability set is broader out of the box — multi-turn conversational editing with retained context is a real feature that the open family does not advertise — and you have no control over any of it.

• Cost basis — Hy Image3.5 preview is $0.024 per 2K image metered on delivered output; LLaDA-Image is free weights plus whatever your GPU costs to run.

• What you get — Hy Image3.5 preview is a hosted API with multi-turn editing and five reference images; LLaDA-Image is four checkpoints (Base and Turbo, BF16 and FP8) and a training recipe.

• Steps to an image — Hy Image3.5 preview is a single call with no step parameter exposed; LLaDA-Image is 50 steps on Base or 2 to 4 on Turbo, which you set yourself.

• Weights — Hy Image3.5 preview publishes none; LLaDA-Image publishes the full family.

• Licence — Hy Image3.5 preview is a commercial service governed by Tencent's terms; LLaDA-Image carries an apache-2.0 tag that our own earlier reporting did not see.

• Evidence — Hy Image3.5 preview has a vendor GSD blind test and no independent Elo; LLaDA-Image has an author-reported Qwen-Image-Bench figure of 53.53 in English and 53.38 in Chinese, and also no independent Elo.

The things only one of these can do

Start with what the weights buy you, because it is more than a licensing preference. A checkpoint is an asset: it cannot be deprecated, repriced, rate-limited or shut off, and a pipeline pinned to a specific file will still run in three years. It can be fine-tuned on your own material, which for a team with a house style or a specific character set is the difference between prompting someone else's model and training your own. It runs offline, which matters if the material is client work under a confidentiality term. And its throughput is a function of the hardware you bought rather than a queue you joined.

Against that, an endpoint buys you an absence of work. Nobody on your team learns a serving stack, nobody sizes a GPU, nobody discovers in production that the FP8 checkpoint needs a different runtime than the BF16 one, and nobody is on the hook when the job fails at 3am. The Tencent model also carries a capability the open family does not claim: multi-turn editing that retains the conversation, which is precisely the mechanism behind holding a character across a long series of edits. LLaDA-Image's editing support is real — it is a unified generation-and-editing family — but conversational state across turns is not something the card advertises.

The honest framing is that these are not competing on quality at all. They are competing on who carries the operational risk, and the two answers are "Tencent" and "you".

What has actually been measured, on either side

Neither model has an independent placement. Artificial Analysis's text-to-image board, read on 23 September 2026, carries no Hy Image3.5 preview row and no LLaDA-Image row. Where the board does have Tencent is the previous generation: HunyuanImage 3.0 Instruct at 964 Elo and HunyuanImage 3.0 at 945, ranked 65th and 70th of 156 models — which is the only third-party measurement of anything in this family, and it is a generation old.

What each side offers instead is its own work. Tencent's is the GSD-style blind preference test run with several hundred of the company's own professional designers, reporting roughly thirty percent over Hy Image3.0, parity with Seedream 5.0 Pro, and a slight edge over Nano-Banana Pro and Qwen-Image-3.0-Pro. Vendor-run, vendor-populated, prompt set unpublished — a real method with a real conflict of interest, and both halves of that sentence belong in the same breath.

LLaDA-Image's is a Qwen-Image-Bench score of 53.53 in English and 53.38 in Chinese, which is an author-reported number on a benchmark the authors chose. It is not more independent than Tencent's test; it is differently unaudited. The card also lists five community Spaces and a monthly download count around a thousand, which tells you the open family has real but modest traction — it is not a model the field has already adopted, and anyone who tells you otherwise has not looked at the number.

A single-panel card titled 'Hy Image3.5 preview vs LLaDA-Image - two unaudited claims and one empty column', subtitled 'Neither model has a third-party placement; both sides are reporting their own work'. Five rows: 'Independent board - no Hy Image3.5 preview row, no LLaDA-Image row, on the 156-model Artificial Analysis text-to-image board'; 'Tencent's own test - GSD-style blind preference with several hundred of its own designers, roughly +30% over Hy Image3.0 and parity with Seedream 5.0 Pro'; 'LLaDA-Image's own numbers - Qwen-Image-Bench 53.53 English and 53.38 Chinese, author-reported on a benchmark the authors chose'; 'Traction for the open family - about 1,021 downloads in the last month and 5 community Spaces on the model card'; 'Previous Tencent generation - HunyuanImage 3.0 Instruct 964 Elo at rank 65 and HunyuanImage 3.0 945 Elo at rank 70, the only third-party measurement in the family'. A footer reads: 'Tencent and inclusionAI figures are vendor-reported and unreproduced by us. Board figures per Artificial Analysis, read 23 September 2026.' The OrcaRouter logo is composited in the bottom-right corner.

Where the money actually goes, at volume

Run the rent side honestly, because it is the only side with a published rate. A hundred thousand 2K images a month at $0.024 is $2,400. Half a million a month is $12,000, or about $144,000 a year, and at that scale a self-hosted 7B-class image model stops being a hobbyist option and starts being an obvious line item. Below roughly ten thousand images a month the arithmetic does not work that way at all: $240 against the cost of a GPU, the power, the storage, the serving stack and somebody's attention is not a close call, and the metered endpoint wins on every axis including the one you are not counting.

The crossover is not a number anyone can hand you, because the self-hosted side depends on hardware you already own, utilisation you actually achieve, and whether the GPU is doing something else when it is not generating images. What can be said precisely is the shape of the curve: the metered model is linear in volume and the self-hosted model is a step function, so the question is not which is cheaper but where your volume sits relative to the step.

One endpoint, whichever way you go

OrcaRouter routes neither of these. We host no Tencent image model — the only Tencent entries in the catalogue are the text-only Hy3 line — and nothing from inclusionAI at all, so Hy Image3.5 preview means Tencent's own cloud and the first-party Tencent products, and LLaDA-Image means the published weights and whatever runtime you point them at. Saying that plainly is more useful than a model page that leaves it ambiguous.

What a routing layer is good for in this decision is the third option it makes cheap. If you are weighing $2,400 a month against a GPU purchase, the useful intermediate is knowing what the same workload costs across models you can reach today without a contract — openai/gpt-image-2, google/gemini-3.1-flash-image-preview, the Imagen 4 family and grok/grok-imagine-image all sit behind one OpenAI-compatible endpoint at provider list price with 0% markup, so a vendor price change is live on our side the same day, and automatic failover means one provider's bad afternoon is not your outage. That is not a substitute for either model in this comparison. It is what stops the comparison from being a two-way choice made in the dark.

A single-panel card titled 'The crossover is not a number anyone can hand you', subtitled 'Metered pricing is the only side of this comparison with a published rate card'. Six rows: '10,000 images a month - about $240 metered, against a GPU, its power, its storage and somebody's attention'; '100,000 images a month - about $2,400 metered'; '500,000 images a month - about $12,000 a month, roughly $144,000 a year'; 'The shape of the curve - the metered model is linear in volume, the self-hosted model is a step function'; 'The question - not which is cheaper, but where your volume sits relative to the step'; 'Before 7 October 2026 - run your own prompt set through the free window on Miora, WorkRally and OnSolo and write down the accept rate'. A footer reads: 'Hy Image3.5 preview pricing per Tencent ($0.024 per 2K image, billed on delivered output). OrcaRouter routes neither model; we host no Tencent or inclusionAI image model.' The OrcaRouter logo is composited in the bottom-right corner.

The one question that decides it

Ask whether your workload has a shape that the weights can capture and the endpoint cannot. If you are generating against a house style, a fixed character set, or a client's own material, and you have the volume and the hardware to make a 7B-class model earn its keep, then LLaDA-Image is the more durable asset and the Apache-2.0 tag — check it at download time, given the history above — is what makes it usable. You are buying optionality and paying for it in operations.

If your workload is bursty, if nobody on the team wants to own a serving stack, if multi-turn editing with retained context is the feature you actually need, or if your volume is under ten thousand images a month, then $0.024 a delivered 2K image is a good price for someone else's problem and the preview label is the main thing to keep in view. The one thing worth doing before 7 October 2026 — while Miora, WorkRally and OnSolo are running the free window — is taking your own prompt set through the tencent model and writing down the accept rate, because that number, not the thirty percent, is the one that decides whether the metered side is worth its meter.