A generated title card headed "AesCode-8B vs Qwen3-8B" with two rounded cards side by side. Left card AesCode-8B carries a browser-window icon rendering a poster page and the lines "Base: Qwen3-VL-8B-Instruct" and "Emits a rendered page"; right card Qwen3-8B carries a document icon with a speech bubble and the lines "Qwen's own generalist" and "Text, 119 languages". A divider between them reads "cousins, not parent and child" and a caption strip across the top reads "AesCode-8B is not a fine-tune of Qwen3-8B". The OrcaRouter logo is composited bottom-right.
Guides & Insights

AesCode-8B vs Qwen3-8B: They Look Like Parent and Child, and They Are Not

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The names invite a mistake. AesCode-8B is an 8B model with a Qwen3-VL-8B-Instruct backbone, Qwen3-8B is the vendor's own 8B model, and the obvious reading is that the first is a fine-tune of the second. It is not. AesCode-8B's published config declares Qwen3-VL-8B-Instruct as the base — the vision-language sibling, a different checkpoint with a different job — and the Hugging Face metadata lists it as a finetune of that model, not of the text-only Qwen3-8B. The two AesCode models in this weight class are cousins rather than parent and child, and that distinction is the whole reason this page is useful: it tells you what Microsoft bought when it started from a vision-language checkpoint, what that cost on the text side, and why a comparison against the family's plain 8B generalist ends up measuring a fork rather than an upgrade.

Both are Apache 2.0 and both are downloadable, but their publication records could not be more different. Qwen3-8B shipped in April 2025 as part of the vendor's Qwen3-generation line, with documentation, ecosystem integration and eighteen months of other people's deployment behind it. AesCode-8B carries no release date anywhere in its files. Microsoft created the Hugging Face repository on 2026-09-29, committed the weights under the message "Release AesCode-8B" at 03:35 UTC on 2026-10-07, and put the training code on GitHub the following day. There is no Microsoft Research post, no release note, no product page and no paper — the model card's citation reads "Under review" with the placeholder year 2027, and the repository showed two downloads as of this writing. Everything quantitative about AesCode-8B below is Microsoft's own and unreproduced by anyone.

The family tree the names hide

The Qwen3-generation line contains several 8B-class checkpoints that share a lineage and not much else. Qwen3-8B is the text-only generalist: roughly 8.2 billion total parameters with about 7 billion non-embedding, grouped-query attention, a 32K-token native context extensible to 131K through YaRN, and training across 119 languages and dialects. It reasons, chats, codes and calls tools, and it is one of the most widely self-hosted 8B models in the open ecosystem for exactly those reasons. Qwen3-VL-8B-Instruct is the multimodal member: the same family generation, a dedicated vision tower on top, image and text in, text out.

Microsoft started from the second one. The AesCode-8B config reads Qwen3VLForConditionalGeneration, with 36 hidden layers, hidden size 4,096, 32 attention heads and 8 key-value heads, a 151,936-token vocabulary and interleaved mrope positions, and it needs transformers 4.57 or newer. The shards total 17,543,339,408 bytes, about 8.8 billion bf16 parameters, which is why Hugging Face's rounded size field says 9B while the vendor says 8B. That is rounding rather than a discrepancy, and the parameter count is not the fork that matters — the input modality and the 24,576-token serving context are.

So the honest framing is not "generalist versus its fine-tune." It is: a text generalist with a long context and eighteen months of production history, against a multimodal research checkpoint that emits a document and has never been independently evaluated.

A generated two-column scoreboard titled "AesCode-8B vs Qwen3-8B - the scoreboard". Left column AesCode-8B reads Base Qwen3-VL-8B-Instruct, Parameters 8.8B, Output HTML and CSS page, Context 24,576 tokens, Languages English only, Evidence vendor rubric and unreproduced. Right column Qwen3-8B reads Base Qwen3-8B itself, Parameters 8.2B, Output text and tool calls, Context 32K native and 131K extended, Languages 119, Evidence 18 months independent. A footer line reads "AesCode-8B figures are Microsoft-reported and unreproduced; Qwen3-8B specifications are Alibaba's." The OrcaRouter logo is composited bottom-right.

What Microsoft spent the training budget on

The design brief is quoted on the model card and it explains the fork: code models "cannot see how layout, hierarchy, and color come together on the canvas," while image generators "compose visually compelling pages but often misrender text, numbers, and logical relationships." AesCode is the attempt to take composition sense from one and verifiability from the other, and the recipe is unusual in a specific way.

Every training prompt is paired with a reference image generated from that same prompt — Microsoft's own runs used GPT-Image-2 for the image and GPT-5.5 to expand a short brief into a detailed content prompt. The reference carries layout and style only; the text prompt stays authoritative on content. Then, at inference, the model barely needs it: withhold the reference from AesCode-8B and its Visual score falls 1.00 point out of a hundred, against 19.55 for the backbone and 10.04 for GPT-5.5. A related result is that reference-based rewards raised prompt-only Visual from 25.17 to 69.71, so the visual prior moved into the weights instead of staying in the input.

The rest is a reward-design story. Supervision is a "design graph" of nodes and edges — text, charts, tables, cards, regions, image slots, with containment, alignment, order and connection as the edges — chosen so every property on the canvas belongs to a named element and can therefore be scored on its own. Seven channels score the result: six deterministic verifiers for execution, text, boundaries, table and chart data, semantic layout and whitespace, plus one sample-specific Visual Graph Rubric judged by a vision-language model. Candidates are rendered in a sandboxed Playwright browser with external requests blocked, and the harness reads back the DOM, computed styles, bounding boxes and a screenshot, which is why tables must be HTML tables and charts must be ECharts specs. Training was cold-start supervised fine-tuning on 3,000 demonstrations, then GDPO over 7,408 prompts for 400 steps on a single node of eight NVIDIA B200s, with each reward channel normalised within its rollout group so a dense rule-based signal cannot drown the sparse visual one. Microsoft measures that normalisation at 5.12 Visual points over plain scalar GRPO.

None of that machinery exists to make AesCode-8B a better general assistant. It exists to make the output checkable, and it is why the model's context is what it is.

Side by side, with the numbers labelled

• Base — AesCode-8B: Qwen3-VL-8B-Instruct, a vision-language checkpoint. Qwen3-8B: the text-only generalist of the same generation.

• Parameters — AesCode-8B: about 8.8B bf16 across four shards. Qwen3-8B: about 8.2B total, roughly 7B non-embedding.

• Inputs — AesCode-8B: text plus an optional reference image. Qwen3-8B: text.

• Output — AesCode-8B: a complete HTML and CSS document. Qwen3-8B: text, tool calls, code.

• Context — AesCode-8B: 24,576 tokens in the reported configuration. Qwen3-8B: 32K native, 131K extended via YaRN.

• Languages — AesCode-8B: the card's tag is "en" alone. Qwen3-8B: 119 languages and dialects.

• Evidence — AesCode-8B: Microsoft's own 300-sample infographic rubric, three generations per prompt, no outside reproduction. Qwen3-8B: eighteen months of independent benchmarks, quantisation work and production deployment.

• Licence — Apache 2.0 both, ungated. A tie, and not the differentiator people assume it is.

The evidence gap is the real comparison

Microsoft reports AesCode-8B at 82.94 Overall on its rubric, ahead of reference-conditioned GPT-5.5 at 81.28 and Claude Opus 4.8 at 80.39, and 31.1 Visual points clear of its own reference-conditioned backbone. Three numbers in that table are worth more than the headline. The first is Style, where AesCode-8B sits at 53.21 and no model in the comparison clears 60 — and Microsoft's definition of Style is design that needs no further visual revision before delivery, so on the one dimension that asks whether a human would ship the page unedited, the model is not the leader. The second is the boundary failure: a severe canvas overflow recurs on 4.3% of the 300 samples, against 34.7% for GPT-5.5. The third is that the visual half of the score comes from an unnamed vision-language judge, which means the deterministic half is reproducible by an outsider and the other half is not.

Qwen3-8B's numbers come from somewhere else entirely, and that matters more than which set is higher. Eighteen months of independent evaluation, community quantisations, serving benchmarks and production deployments is a different kind of evidence: it is not a claim, it is a track record. You can find out how Qwen3-8B behaves under four-bit quantisation on a specific card because someone has published it. You cannot do that for AesCode-8B, because as of this week nobody outside Microsoft has run it on anything.

And there is no shared metric between the two tables. Qwen3-8B has no infographic layout score; AesCode-8B has no published general-reasoning benchmark at all. The comparison is not close or not close — it is unavailable.

What running each one actually takes

A screenshot of the OrcaRouter model page for qwen/qwen3-vl-8b-instruct showing the Vision, Tools and JSON capability badges, the byline "by Qwen - 2025-10-14", the description "Qwen3-VL 8B Instruct - open-weight small vision-language model, 8B params, 128k context, no thinking mode", the pricing row reading $0.18 input and $0.70 output per million tokens, and the OpenAI-compatible code sample against the https://api.orcarouter.ai/v1 base URL.

Qwen3-8B is the easy one. An 8.2B text model quantises onto a single consumer card, has eighteen months of tooling behind it across every serving framework and local runtime, and behaves predictably. The context extension to 131K is documented rather than folklore, and the 119-language coverage is the reason it turns up in so many multilingual pipelines.

AesCode-8B is a serving project. The card's own command is vllm serve microsoft/AesCode-8B --limit-mm-per-prompt image=2 --max-model-len 24576, with transformers 4.57 or newer for the non-vLLM path. The 17.5 GB of bf16 weights have to fit beside a KV cache for 24,576 tokens and up to two images, so a single 24 GB card is technically sufficient and uncomfortably tight, and 40-48 GB or two 24 GB cards is the realistic floor. Budget a renderer too, because every claim about this model's quality was made by rendering its HTML output in a sandboxed browser — if you want to know whether your own outputs are good, that is the harness you have to stand up. And note that the 24,576-token context was exercised on single infographic pages, not the multi-slide decks the model is pitched at, so context pressure on a real deck is untested territory.

Availability is the other half of the same point. AesCode-8B is not something OrcaRouter routes, and we could not find it served anywhere else either; an evaluation today means weights on your own hardware. Its ancestor is a different story — Qwen3-VL-8B-Instruct is callable through OrcaRouter right now at $0.18 per million input tokens and $0.70 per million output over a 131,072-token context, which makes it the cheapest way to see what the untuned version of this exact architecture does on your own infographic prompts. Further up the same line, Qwen3.8-27B is routable at $0.33 and $2.40. Provider list price passes through with no markup added and failover between providers is automatic, so prototyping the prompt against the hosted backbone costs one key rather than a GPU-week, and only the prompts that actually render well earn the reproduction work.

A Hugging Face screenshot of the microsoft/AesCode-8B model card showing 0 likes at capture time, the Microsoft organisation follower count, the Image-Text-to-Text, Transformers, Safetensors and English tags with qwen3_vl and code-generation, an Apache-2.0 licence badge, and the opening card text describing AesCode as generating information-rich visual artifacts such as slides, posters and dashboards as HTML/CSS, followed by the line that AesCode-8B starts from Qwen3-VL-8B-Instruct.

Which one you want depends on the artifact

Pick AesCode-8B when the deliverable is a page and it has to be editable. Structured HTML output that diffs in Git, tables as real tables and charts as ECharts specs, is a genuinely different artifact from a paragraph of text, and no amount of general capability substitutes for it. Take the trade knowingly: an unannounced research checkpoint with a double-digit download count, a 24K context, English only, no independent evaluation, a Style ceiling its own vendor publishes, and a serving bill measured in tens of gigabytes.

Pick Qwen3-8B for everything else — which, for most teams, is everything. It is the proven generalist at this weight, it has a longer context, it speaks a hundred and nineteen languages, it runs on hardware you already own, and it has eighteen months of other people's production experience behind the decisions you are about to make. If you have no specific need to emit rendered pages, this comparison has one answer.

The case for wanting both is real, and it is not a tie-break. A pipeline that reads documents and produces decks uses two models with two genuinely different output types, and the useful posture is to keep the page renderer on hardware that can hold it and route the reading and reasoning work to something hosted behind the same endpoint as everything else.

What would make this a closer page

Three things, and only one of them is about benchmarks. An independent reproduction of the 82.94 and the 1.00-point reference drop would move AesCode-8B from vendor claim to result, and would let the 8B-versus-8B size question be answered on merit. A provider picking the checkpoint up would collapse the deployment column and turn "a GPU-week" into "an API call." And the paper the card cites as under review — "AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards" — would answer the question the model card cannot: what the visual rubric does when the design graph it scores against is itself wrong.

Until one of those lands, the defensible reading is the narrow one. AesCode-8B is the more interesting of these two models and the less usable one. It is very good at the parts of visual design a machine can check, middling at the part that cannot be automated, and it belongs to the same Qwen3-generation family as the model it is being compared to without being descended from it. Qwen3-8B is the one you can put in a pipeline this afternoon.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily