A generated title card headed "AesCode-8B vs Gemma 4 12B" with two rounded cards side by side. Left card AesCode-8B carries a browser-window icon and the lines "Microsoft research checkpoint" and "Emits a rendered HTML page"; right card Gemma 4 12B carries a magnifier-over-document icon and the lines "Google DeepMind, launched 2026-06-03" and "Answers questions about images". A divider between them reads "one renders, one replies" and a caption strip across the top reads "unannounced release - weights only, no hosted endpoint". The OrcaRouter logo is composited bottom-right.
Guides & Insights

AesCode-8B vs Gemma 4 12B: Two Small Vision Models, Two Completely Different Outputs

Author

Magnus Corvin

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

AesCode-8B and Gemma 4 12B both take an image and a sentence and both are Apache-2.0 checkpoints you download rather than rent, which is why search engines will happily put them side by side. The resemblance is real and it is also shallow. Gemma 4 12B returns an answer — a description, a transcription, a code snippet, a reasoning trace. AesCode-8B returns a rendered page: a slide, a poster or a dashboard as a complete HTML and CSS document you can open in a browser and edit by hand. Same input shape, same rough weight class, and an output that shares almost nothing with the other. That single difference decides nearly every practical question below, including which one you can put in front of a user tomorrow.

Worth stating up front, because it shapes how much weight the numbers can carry: Gemma 4 12B was launched by DeepMind on 2026-06-03 with a model card, documentation and an ecosystem, and it has been independently tested since. AesCode-8B was not launched at all. Microsoft's repository for it was created on 2026-09-29 and the weights landed in a commit titled "Release AesCode-8B" at 03:35 UTC on 2026-10-07, with the training and evaluation code appearing on GitHub the following day. There is no Microsoft Research post, no product page, no release note and no preprint — the model card's citation reads "Under review" with the placeholder year 2027. As of this writing the 8B repository shows two downloads and one like. Everything quantitative about AesCode-8B is therefore Microsoft's own, and nothing has been reproduced by anyone else.

The two models, on their own terms

Gemma 4 12B is a mid-size open model, an 11.95-billion-parameter dense checkpoint whose defining design choice is that it has no separate vision or audio encoder: image patches and audio waveforms go straight into the transformer. It carries a 256K-token context and wants about 16 GB of VRAM, with the BF16 weights around 27 GB and quantised builds below 7 GB. The vendor's own card — vendor-reported, and labelled as such here — lists DocVQA 94.9, InfoVQA 88.4, MMLU-Pro 77.2, GPQA Diamond 78.8, AIME 2026 77.5 and LiveCodeBench v6 72.0 in thinking mode. Independent evaluation is the part that matters more: Artificial Analysis scores the reasoning configuration at an Intelligence Index of 14, measures roughly 114 tokens per second, and prices a typical served call near $0.10 per million input tokens and $0.30 per million output. Three and a half months of deployment sit behind it.

AesCode-8B is a fine-tune of Qwen3-VL-8B-Instruct, published with four safetensors shards totalling 17,543,339,408 bytes — about 8.8 billion bf16 parameters, which is why Hugging Face's rounded size field reads 9B while the vendor says 8B. The published config is a straight Qwen3-VL recipe: 36 hidden layers, hidden size 4,096, 32 attention heads with 8 key-value heads, a 151,936-token vocabulary, and a requirement for transformers 4.57 or newer. Microsoft's reported runs used a 24,576-token context with up to 12,000 output tokens, temperature 0.8, top-p 0.95 and three generations per prompt with no selection. Licence is Apache 2.0, ungated, no acceptable-use addendum. What is not in the box is the training and evaluation data, and any hosted endpoint — we could not find AesCode-8B served anywhere, and it is not in our own catalogue.

What the models were actually trained to do

The design brief is quoted in the model card and it is worth reading as the whole thesis: code models "cannot see how layout, hierarchy, and color come together on the canvas," while image generators "compose visually compelling pages but often misrender text, numbers, and logical relationships." AesCode is the attempt to keep the composition sense and the verifiability at once.

The mechanism is unusual in one specific way. Every training prompt is paired with a reference image generated from that same prompt, which supplies layout and style while the text prompt remains the authority on content. Then, at inference, the model barely needs the picture: withhold the reference from AesCode-8B and its Visual score drops 1.00 point out of a hundred. Withhold it from Qwen3-VL-8B-Instruct and the drop is 19.55. Withhold it from GPT-5.5, evaluated under the same harness, and it is 10.04. The visual prior ended up in the weights rather than in the input, and that is the genuine research result underneath the headline.

Gemma 4 12B's training target is the ordinary multimodal one and it is not a criticism to say so: read the image, answer the question, follow the instruction, call the tool. Nothing in its card suggests it was ever pointed at producing a rendered artifact, and no published Gemma 4 12B evaluation measures document layout quality, canvas overflow or design consistency, because those are not the model's metrics.

The scoreboard, and whose rubric it is

A generated two-column scoreboard titled "AesCode-8B vs Gemma 4 12B - the scoreboard". Left column AesCode-8B reads Parameters 8.8B dense, Inputs text plus one image, Output a rendered HTML page, Context 24,576 tokens, Evidence Microsoft rubric with no outside test, Hosted nowhere we can find. Right column Gemma 4 12B reads Parameters 11.95B dense, Inputs text, image, audio, video, Output text and tool calls, Context 256K tokens, Evidence AA Intelligence Index 14, Hosted cheap served calls. A footer line reads "AesCode-8B figures are Microsoft-reported and unreproduced; Gemma 4 12B figures per Artificial Analysis." The OrcaRouter logo is composited bottom-right.

• Parameters — AesCode-8B: 8.8B bf16, dense. Gemma 4 12B: 11.95B dense.

• Inputs — AesCode-8B: text plus an optional reference image. Gemma 4 12B: text, image, audio and video.

• Output — AesCode-8B: a complete HTML and CSS document. Gemma 4 12B: text, including tool calls.

• Context — AesCode-8B: 24,576 tokens in the reported configuration. Gemma 4 12B: 256K tokens.

• Licence — both Apache 2.0, ungated, weights included.

• Evidence — AesCode-8B: Microsoft's own 300-sample infographic rubric, three generations per prompt, no independent reproduction. Gemma 4 12B: a vendor card plus an independent Intelligence Index of 14 and measured throughput on Artificial Analysis.

Now the part the scoreboard cannot carry. Microsoft reports AesCode-8B at 82.94 Overall on that 300-sample set, ahead of reference-conditioned GPT-5.5 at 81.28 and Claude Opus 4.8 at 80.39, and 31.1 Visual points clear of its own Qwen3-VL-8B backbone. Read all of it as vendor-reported and unreproduced, on a rubric the vendor built, on samples the vendor chose, scored by a harness the vendor trained the model against. The card is honest enough to publish a dimension no model in the comparison clears 60 on — Style, where AesCode-8B sits at 53.21 and GPT-5.5 leads at 57.91 — and Microsoft's own definition of that dimension is design that needs no further visual revision before delivery. On the one axis that asks whether a human would ship the page unedited, the 8B model is not the leader. That is the number to hold onto when someone quotes the 82.94.

Where they genuinely overlap

There is exactly one shared shelf, and it is narrower than it looks. If you are evaluating small open vision models for a document-heavy pipeline, both are candidates — one for reading a document, the other for producing one. A team that generates internal dashboards as HTML and also needs to extract figures out of uploaded PDFs touches both products in the same week.

The split inside that overlap is clean. Gemma 4 12B is the model you point at an incoming document. AesCode-8B is the model you point at a brief when the deliverable is a page. Neither substitutes for the other, and a pipeline that wants both is a two-model pipeline rather than a choice.

A Hugging Face screenshot of the microsoft/AesCode-32B model card showing tags for Image-Text-to-Text, Transformers, Safetensors, English and qwen3_vl, an Apache-2.0 licence badge, 1 like from the Microsoft organisation, and the card opening text that AesCode generates information-rich visual artifacts such as slides, posters and dashboards as HTML/CSS with output that stays structured, editable and verifiable.

Deployment, which is where the asymmetry actually bites

Gemma 4 12B's serving story is finished. You download it, quantise it if you want, run it on 16 GB of VRAM, and there is a large base of tooling that already does this. If you would rather not run it at all, a served call is cheap and independently priced.

AesCode-8B's serving story is a command line and some arithmetic. The model card gives the route directly — vllm serve microsoft/AesCode-8B --limit-mm-per-prompt image=2 --max-model-len 24576, with transformers 4.57 or newer for the non-vLLM path. 17.5 GB of bf16 weights have to sit alongside a KV cache for 24,576 tokens and up to two images, so a single 24 GB card is technically enough and uncomfortably tight; 40-48 GB, or two 24 GB cards, is the realistic floor. Budget a renderer as well, because every quality claim about this model is made by rendering its HTML output in a sandboxed browser, so scoring your own results means standing up something Playwright-shaped with external requests blocked.

Then there is the honest state of availability. We do not route AesCode-8B, and neither, as far as we can find, does anyone else; an evaluation today means weights on your own hardware. The nearest thing that is callable is the architecture the model was fine-tuned from. Qwen3-VL-8B-Instruct is routable through OrcaRouter now at $0.18 per million input tokens and $0.70 per million output over a 131,072-token context, and the two checkpoints we do carry are priced the same way — Gemma 4 26B-A4B at $0.06 and $0.33, Gemma 4 31B at $0.13 and $0.38. Provider list price passes through with no markup added, and failover between providers happens automatically, so the cheap way to find out whether your prompts render well at all is to test the untuned backbone first and spend the GPU-week only on the prompts that survive.

A screenshot of the OrcaRouter model page for google/gemma-4-31b-it, showing the Featured badge, the byline "by Google - 2026-04-02", a description of Gemma 4 31B Instruct as a 30.7B dense multimodal model with text and image input and a 256K-token context window, and the pricing row reading $0.13 input and $0.38 output per million tokens with measured latency figures.

Which one you actually want

Pick AesCode-8B if the deliverable is a page. The model emits structured, editable HTML — tables as HTML tables, charts as ECharts specs — which means the output diffs in Git and can be restyled by a designer without regenerating it. Accept the trade: an unannounced research checkpoint with a double-digit download count, no independent evaluation, no hosted endpoint, a 24K context that has only been validated on single infographic pages, and an English-only language tag on the card.

Pick Gemma 4 12B for everything else. It reads documents better than most models twice its size, it handles audio, it follows tool instructions, it has a quarter of a million tokens of context, and it has three and a half months of other people's experience behind it. If you are choosing one of these two to be your small vision model, and you have not specifically been asked to produce slide decks as code, this is the answer.

And if the honest requirement is both, do not treat it as a tie-break. Route the reading work to a hosted model and keep the rendering work on whatever machine can hold the weights.

What would change this page

Three signals. An independent reproduction of the 82.94 and the 1.00-point reference drop would move AesCode-8B from a vendor claim to a result, and would make the 8B-versus-12B size question genuinely interesting rather than nominally so. An inference provider picking the checkpoint up would collapse the deployment column, which is currently the entire practical difference. And the paper Microsoft cites as under review — "AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards" — would answer the questions the model card cannot, starting with how the visual rubric behaves when the design graph itself is wrong. Until one of those lands, the defensible reading is narrow and real: AesCode-8B is very good at the parts of visual design a machine can check, middling at the part it cannot, and Gemma 4 12B is a finished product doing a different job.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily