A hero title card for 'AesCode 32B vs Qwen3.8 27B' subtitled 'the fine-tune and the ancestor', drawn as a lineage diagram: a Qwen3-VL-32B-Instruct checkpoint card on the left feeds an arrow into an 'AesCode 32B fine-tune' card labelled 'self-host only, 65 GB', while a separate newer card below-right labelled 'Qwen3.8 27B' carries a live API marker reading 'callable today'; the OrcaRouter logo sits in the bottom-right corner.
Guides & Insights

AesCode 32B vs Qwen3.8 27B: Microsoft Built It On Qwen, and One Of Them Is Callable

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

The detail that reorganises this whole comparison is on the second line of the model card: AesCode 32B starts from Qwen3-VL-32B-Instruct. Microsoft's design-specific code model is a Qwen3-VL fine-tune, Apache 2.0, sitting on a Hugging Face repository created September 29, 2026, with the commit titled "Release AesCode-32B" dated October 7, 2026 and no announcement anywhere from Microsoft. Qwen3.8 27B is the other end of that lineage: the vendor's own open-weight 27B dense multimodal model, released August 14, 2026 under Apache 2.0, the successor architecture to the VL checkpoint Microsoft fine-tuned. So this is not a matchup between two vendors' flagship efforts. It is a matchup between one lab's general-purpose open-weights model and a competitor's research fine-tune of the family, and the practical difference between them turns out to be almost entirely about what you are allowed to do with the file once you have it.

Same licence, same family, different amount of model

Both are Apache 2.0, both take images as well as text, and both return text. Everything after that diverges, and the deployment row diverges hardest.

• Parameters — AesCode 32B: 33B dense on a Qwen3-VL-32B-Instruct base. Qwen3.8 27B: 27B dense, 28B counting the vision encoder, 64 layers at a hidden size of 5120.

• Memory to run it — AesCode 32B: the card advises planning roughly 65 GB of accelerator memory for the bf16 parameters alone, and its own serving example uses four-way tensor parallelism. Qwen3.8 27B: about 55.6 GB in bf16.

• Context — AesCode 32B: 24,576 tokens in the reported configuration, with training-time prompts and responses capped at 8,192. Qwen3.8 27B: 262,144 tokens native, extensible to 1,000,000 with YaRN scaling.

• Attention — AesCode 32B: inherited from the Qwen3-VL backbone. Qwen3.8 27B: hybrid, 48 Gated DeltaNet linear-attention layers to 16 full-attention layers at a 3:1 ratio, which is what makes a 262K window affordable to serve.

• What it is for — AesCode 32B: generating information-rich visual artifacts as editable HTML and CSS, plus research on multimodal code generation. Qwen3.8 27B: general agentic and multimodal work — coding, computer use, document and video understanding, tool calling.

• Where you get it — AesCode 32B: a Hugging Face download; no inference provider deploys it. Qwen3.8 27B: weights, or a hosted endpoint.

Twice the context, slightly smaller footprint, one generation newer backbone, and a licence that is identical. The only row where AesCode 32B wins outright is the first, and it wins it on a task-specific fine-tune.

What each was actually measured on

The two evaluation regimes do not touch, and pretending otherwise is the standard error on pages like this one.

Qwen3.8 27B's card is broad and specific at once. Coding: Terminal-Bench 2.1 at 73.0, SWE-bench Pro at 61.7, QwenSWEBench at 79.0, LiveCodeBench v6 at 90.3. Reasoning: GPQA Diamond 89.2, Humanity's Last Exam 30.8. Computer use: OSWorld-Verified 84.3, WebArena-Verified 64.8, AndroidWorld 81.9. Vision: MathVision 94.6 with the vendor's inference-time enhancement and 90.0 without it, OmniDocBench 1.5 at 91.1. All of it the vendor's own numbers on the vendor's own harness — with a footnote worth reading, that the SWE-bench Pro and QwenSWEBench runs used the Claude Code harness, so they are not directly comparable with models scored on a different one.

AesCode 32B has one benchmark and it is a single-purpose one: 300 infographic samples, three generations per prompt with no selection, rendered and scored across seven channels. The headline is an Overall of 85.89 with the reference image supplied, against 81.28 for GPT-5.5 and 80.39 for Claude Opus 4.8 under the same harness — plus the claim that AesCode 32B beats its own Qwen3-VL-32B backbone by 22.4 points on the Visual dimension and 24.8 on Overall.

Put those next to each other and nothing lines up. Terminal-Bench measures whether a shell task completes; the AesCode evaluation measures whether an infographic's text is legible, its layout does not overflow the canvas and its table is a real HTML table. There is no shared metric, no shared harness and no shared task. The only honest statement is that AesCode 32B is much better at design artifacts than its own backbone, and Qwen3.8 27B is a strong general model — and neither claim tells you anything about the other.

A two-column comparison scoreboard for AesCode 32B and Qwen3.8 27B: the AesCode 32B column reads 'Parameters: 33B dense (Qwen3-VL-32B base)', 'Memory to run it: about 65 GB bf16, 4-way TP', 'Context: 24,576 tokens', 'Attention: inherited from the Qwen3-VL backbone', 'Where you get it: a Hugging Face download', 'Measured on: one 300-sample infographic benchmark'; the Qwen3.8 27B column reads 'Parameters: 27B dense (28B with vision tower)', 'Memory to run it: about 55.6 GB bf16', 'Context: 262,144 native, 1M with YaRN', 'Attention: hybrid Gated DeltaNet, 3:1 ratio', 'Where you get it: weights or a hosted endpoint', 'Measured on: Terminal-Bench 2.1, SWE-bench Pro, GPQA Diamond and more'; a footer line reads 'AesCode 32B figures Microsoft-reported on its own harness; Qwen3.8 27B figures Alibaba-reported with independent scoring from Artificial Analysis.' and the OrcaRouter logo sits in the bottom-right corner.

The evidence gap is real this time, in one direction

Vendor benchmarks are the norm in this industry, so it matters that these two differ in how much of their vendor claims somebody else has checked.

Qwen3.8 27B shipped with the vendor's own table and then accumulated outside results within weeks. The model's independent record includes a legal-agent study run by the legal-AI company Harvey with memory startup Engram, in which an adapted Qwen3.8 27B led every model the study tested including Claude Opus 4.8 on 250 legal tasks — with the important caveat that the adapted model had studied the test firm's 100M-token corpus first, so it is a result about adaptation as much as about the model. It includes a placement on Code Arena's WebDev leaderboard and the top open-weight spot on Arena.ai's Image-to-WebDev leaderboard, both of which are vendor-relayed ranking claims rather than published methodology. And it includes a third-party board with a published score: Artificial Analysis records Qwen3.8 27B at 33.70 on its Intelligence Index with a cost per finished task around $1.01, and publishes the token breakdown behind that.

AesCode 32B has none of that. No arena placement, no leaderboard listing, no third-party reproduction, no independent throughput test — and not because anyone looked and found it wanting, but because a checkpoint that no provider serves cannot be evaluated by an outside harness in the first place. Its numbers are the vendor's, computed by the same verifier stack that trained the model, on samples the vendor chose, against baselines the vendor ran. The release is unusually transparent about method — the reward channel names, the rollout counts, the decoding parameters and the failure rates are all published — but transparency about how a number was produced is not the same as somebody else producing it.

The one that costs a card on file

This is where the lineage stops being trivia and starts being the decision.

AesCode 32B asks you for 65 GB of accelerator memory, a multimodal serving path, and the willingness to run an unannounced model as an early adopter. Its own serving example reaches for four-way tensor parallelism, and the model is documented for a 24,576-token context with no hosted option to fall back on. If you already own the hardware and your pipeline genuinely needs design-specific fine-tuning, that is a reasonable trade. For most teams it is a two-week project before the first useful output.

Qwen3.8 27B is self-hosted on OrcaRouter's own infrastructure and callable today at $0.33 per million input tokens and $2.40 per million output, with a 262,144-token context and text, image and video input. Those are list rates passed through with no markup added on our side, and it sits on the same key as 200-plus other models, so the comparison against any alternative on the catalogue is a request parameter rather than a second contract. That is the whole of the practical case, and it is a large one: the model that is a generation behind AesCode 32B's backbone in family terms is the one you can evaluate this afternoon, on your own prompts, for the price of a few cents.

There is a second-order point here that the routing angle makes concrete. Because AesCode 32B is a fine-tune of a Qwen3-VL checkpoint, the family gives you a cheap way to test whether the hosted side of the architecture works at all before you commit to self-hosting anything. Qwen3-VL 235B A22B Instruct is available at $0.40 per million input and $1.60 output, and Qwen3-VL 8B Instruct at $0.18 and $0.70. Neither is AesCode 32B and neither has been trained for design quality — but "can a vision-language model read our screenshots and write the markup we need" is answerable with them for cents, and automatic failover underneath the same endpoint means a bad provider day does not take the pipeline down.

A screenshot of the Hugging Face model page for microsoft/AesCode-32B, showing the model card header, the Apache 2.0 licence tag and the model-index table carrying the 300-sample infographic evaluation with the 85.89 Overall score (captured October 11, 2026).

Who should pick which

Take AesCode 32B if the artifact is the deliverable. You produce slides, dashboards, posters or reports in volume, you need structured output that stays editable and diffable — the model emits a complete HTML document, uses real HTML table structures and ECharts specs so both stay inspectable — and you have the hardware or the budget to rent it. Go in accepting three things: no independent verification of any number, roughly 65 GB for the weights, and a 24K context that will constrain long documents. The 4.3% severe canvas-boundary failure rate the card reports, against 34.7% for GPT-5.5, is the most encouraging single figure in the release, and it is also Microsoft's.

Take Qwen3.8 27B if you need a general open-weight multimodal model you can actually run and actually serve. 262K context, extensible to a million; a deployment footprint that fits where AesCode 32B does not; a licensing position no different; some independent measurement of both its quality and its cost per finished task; and a hosted option that means you can start today and self-host later without rewriting the integration. Its staged, layer-wise attention design is the reason the long context is affordable, and long context is the thing design and document pipelines tend to need most.

And if you need both — a design-artifact generator and a general workhorse — the sensible split is not to run two 30B-class vision-language models on your own kit. Route the general traffic to the hosted side and keep the specialist self-hosted, behind one key, where a provider outage on either path is a failover event rather than an incident.

A screenshot of the OrcaRouter model page for Qwen3.8 27B, showing the model name, the 262,144-token context window, the text/image/video input modalities and the self-hosted pricing of $0.33 per million input tokens and $2.40 per million output.

The thing to watch

AesCode 32B's position changes completely if it becomes callable. Right now the model's only real weakness is access, and that is a temporary condition — an inference provider picking up the checkpoint, or Microsoft publishing an endpoint, converts a 65 GB procurement into a model selection. Until then, the honest summary of this matchup is that the fine-tune is better at one job, the general model is better at being usable, and the general model happens to be the ancestor: Microsoft chose a Qwen3-VL checkpoint to build on because it was the strongest open multimodal base available, which is itself the most useful thing this comparison has to say about Qwen3.8 27B.

Compared in this article1

Detected from this article · Benchmarks: Artificial Analysis · updated daily