
AesCode-32B: Microsoft's Quiet 33B Model That Writes Decks as Editable HTML
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 87 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 115 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 47 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 777 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 61 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 452 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
The timestamps on AesCode-32B disagree with each other, and the disagreement is most of what there is to report. The Hugging Face repository microsoft/AesCode-32B was created on 29 September 2026, but everything in it arrived in a single commit dated 7 October 2026 whose message is simply "Release AesCode-32B" — and the training stack that makes the result reproducible was pushed to github.com/microsoft/AesCode on 8 October. Microsoft has announced nothing: no blog post, no arXiv preprint, no leaderboard submission, no model page on its own site. What exists is a 33-billion-parameter vision-language model that takes a prompt and emits a complete, self-contained HTML document — a slide, a poster, a dashboard — plus a smaller companion, microsoft/AesCode-8B. Both are fine-tuned from Qwen3-VL vision models — Qwen3-VL-32B-Instruct and Qwen3-VL-8B-Instruct respectively — both are Apache 2.0, and both are downloadable today.
The repo id reads microsoft/AesCode-32B and the card opens with the trick: AesCode pairs your prompt with an image generated from that same prompt, uses the image as an aesthetic reference, and follows the text for the actual content. Image generators compose a handsome page and misrender the numbers on it; code models get the numbers right and cannot see how the page looks. AesCode is an attempt to have both, and the honest summary of it as of 11 October 2026 is that the weights are real and checkable, the benchmark numbers are the lab's own on a harness the lab wrote, and nobody outside Microsoft has published a number for it yet. This piece keeps those three categories apart all the way down.
What is actually in the repository, byte for byte

Start with what you can verify without trusting a word of the model card. The 32B repository holds fourteen safetensors shards totalling 66,714,912,704 bytes, which at BF16 is roughly 33.4 billion parameters — consistent with the card's "33B params" and its warning that the weights alone need about 65 GB of accelerator memory. The config declares Qwen3VLForConditionalGeneration as the architecture and qwen3_vl as the model type, so this is not a new architecture and does not need one: it loads with transformers>=4.57, and the card ships a vLLM command that serves it across four tensor-parallel ranks with two images allowed per prompt and a 24,576-token ceiling.
The engagement counters are the quietest part of the release. Two downloads and one like on the 32B repository at the time of writing. There is no entry in Hugging Face's inference-provider mapping, meaning no hosted endpoint is wired up behind the repo page, and no GGUF, MLX or llama.cpp build is advertised anywhere. For a model bearing Microsoft's name and carrying results that beat GPT-5.5 on the vendor's own table, that is a strikingly small footprint — and it is the strongest available evidence that this went up without a launch behind it.
The companion code repository fills in the other half of the timeline, and it is where the release date question gets genuinely ambiguous. microsoft/AesCode was created on 23 July 2026 — ten weeks before the weights — and carries fourteen commits, every one of them authored by the same contributor and every one of them timestamped within twenty seconds of the others on 8 October 2026, between 23:14:04 and 23:14:24 UTC. They include the spec for generating prompts and verifiable requirements, the data pipeline that builds design graphs and rubric questions, a cold-start SFT stage, the GDPO reinforcement-learning loop, a Playwright-based rendering verifier, tests, and the README that documents all of it. There are no releases and no tags, the repository description is empty, and it has one star.

So which date is the release? The repository record says 29 September. The commit that carries the model says 7 October. The code that explains how the model was made says 8 October. Three timestamps inside a single two-week window, none of them accompanied by a sentence from Microsoft saying "we are shipping this." Treat 7–8 October as the operative date for the artifacts most people will actually download, and 29 September as the date the repository was reserved. Anyone who tells you AesCode-32B "launched" on a specific day is choosing one of those timestamps on your behalf.
The mechanism: an image as aesthetic guidance, a graph as the reward
The card's technical claim is narrow and specific, which is a point in its favour. The model is trained with cold-start supervised fine-tuning on 3,000 demonstrations at a learning rate of 1e-5, then with GDPO — a group-relative policy-optimisation variant — over 7,408 prompts for 520 steps. The 8B model used the same recipe and stopped at 400 steps. The RL run used verl's FSDP-vLLM hybrid engine with no critic and no separately trained reward model, AdamW at a constant 5e-6 with no warmup, 128 prompts per step with eight rollouts each, and prompt and response each capped at 8,192 tokens.
What makes the reward unusual is that it is not a single scalar. Each training target is described as a design graph covering the whole canvas, so individual properties can be attributed separately. From that graph come seven channels — execution, text, boundary, tablechart, layout, whitespace and design — each normalised within its rollout group before aggregation so that one dominant signal cannot drown out the others. Deterministic verifiers score what can be parsed from the code and its rendering; a vision-language judge scores what cannot, using a rubric tied to the graph's own elements and relations. Candidate HTML is scored by rendering it in a sandboxed Playwright browser with external requests blocked, which exports the DOM, computed styles, bounding boxes, console status and a screenshot.
The engineering consequence is worth stating because it shows up in the output you receive: the model is trained to emit tables as real HTML table structures and charts as ECharts specifications, so both are directly inspectable rather than baked into pixels. That is the difference between a deck you can hand to a designer and a deck you can hand to a linter. The reproduction requirements are correspondingly heavy — the README asks for Python 3.10, CUDA 12.6 and a node of eight B200 GPUs, pins a specific verl commit and states plainly that a patch against it is required because stock verl lacks Qwen3-VL support, and warns that without the Playwright system libraries the browser fails at launch and pages score zero, and that without the pinned OCR stack the corresponding reward channel returns zero instead of abstaining and silently corrupts the signal.
The benchmark table, and the four reasons to hold it loosely
AesCode-32B's headline numbers come from 300 infographic samples, three generations per prompt at temperature 0.8 and top-p 0.95, up to 12,000 output tokens each, with no selection among the generations. Scores are percentages. Reference-conditioned, the 32B model reports Text 95.34, Boundary 97.27, Table/Chart 90.37, and a Rule average of 94.33; on the visual side Content 85.76, Layout 90.58, Style 55.99, for a Visual average of 77.44 and an Overall of 85.89. In the same table, GPT-5.5 with a reference scores 81.28 Overall and Claude Opus 4.8 with a reference scores 80.39, while the Qwen3-VL-32B-Instruct backbone the model was trained from scores 61.10.

Four caveats belong in the same breath as those figures, and none of them is a smear on the work. First, every row including the GPT-5.5 and Claude Opus 4.8 rows was run by Microsoft on Microsoft's harness with Microsoft's rubric — those are not the other labs' numbers, they are Microsoft's measurements of competitors' models, and the card itself calls the rubric "sample-specific" to each design graph. Second, the rubric is produced by the same pipeline that generated the training data, which is exactly the setup in which a benchmark can drift toward a model's strengths; the card is candid about the limits of that rubric — Style, which requires a design to need no further visual revision before delivery, is called "the shared ceiling for every system" and no model in the table clears 60. Third, there is no independent measurement of this model anywhere: no third-party leaderboard entry, no reproduction, and given two downloads, almost certainly nobody outside the lab running it yet. Fourth, the comparison is subtly asymmetric in a way worth noticing — AesCode-32B's rows are all reference-conditioned, so the model is being measured in the configuration it was trained for, which the card acknowledges by showing that quality is still highest when a reference is supplied.
The one result in the table that holds up better than the headline is a robustness claim rather than a quality claim. Withholding the reference image at inference costs AesCode-8B only 1.00 Visual point, against 19.55 for its Qwen3-VL-8B-Instruct backbone and 10.04 for GPT-5.5. The card's argument is that reference-conditioned training internalises visual planning into the policy rather than teaching the model to copy what it sees. That is vendor-reported and unreproduced, and it is also the kind of claim that a single independent run would settle — and the kind that matters most in production, where you will not always have a reference image to hand.
What it costs to run, and what that means for the comparison
Nothing about this model is cheap to self-host. Ten to twelve thousand output tokens is the working size of a single artifact, and a full HTML document with an ECharts spec is closer to the top of that range than the bottom, so every generation is a long decode. The card's own serving recipe asks for four GPUs at tensor parallelism to hold roughly 65 GB of BF16 parameters, and the training recipe asks for eight B200s. That is a real machine, not a hobby deployment, and it sets the terms of the comparison: the models AesCode-32B is being measured against are rented by the token, and the model itself is rented by the GPU-hour, whether you own the hardware or not.
The practical shape of that comparison is the point a routing layer exists for, and it is worth being exact about what we do and do not host. AesCode-32B is not on OrcaRouter's catalogue and we do not serve it — there is no hosted endpoint for it anywhere I can verify, Microsoft's included. What is on the catalogue is the other side of the table: the hosted models you would benchmark a self-hosted artifact generator against, including GPT-5.5 and the smaller Qwen3-VL vision models, reachable through one API key at provider list price with 0% markup, which means a vendor price change is live on our side the same day. Comparing a rented endpoint against a model you run yourself does not need a second contract or a second SDK, and a routing DSL lets a self-hosted call sit beside hosted ones behind a single endpoint. If AesCode-32B turns out to be good at the one thing its card claims, the cost of finding out is a GPU bill, and the cost of the alternatives you are comparing it to is one key you probably already have.
What you can do with it today, and what does not exist
• Download and run it — the weights are Apache 2.0, following the Qwen3-VL backbone, with fourteen BF16 shards and a working transformers path above version 4.57 and a vLLM recipe in the card.
• Reproduce the training — the code is MIT-licensed and complete enough to be meaningful: the reward verifier, the rubric builder, the SFT and GDPO stages, a pinned verl commit plus the patch that adds Qwen3-VL support, and a README that lists the failure modes rather than hiding them.
• Evaluate it reference-free — the model accepts prompts with or without the reference image, and the card's most testable claim lives in that configuration.
• Get an API for it — you cannot. There is no hosted endpoint, no inference-provider mapping on the repository, and no GGUF or MLX build; running this means running the hardware.
• Read the paper — you cannot yet. The card links a paper titled AesCode: Aesthetic Code Generation with Decoupled Cross-Modal Rewards, and its own BibTeX entry gives the venue as "Under review" and the year as 2027. A search of arXiv returns no paper with that title. There is a different, earlier Microsoft paper with a near-identical name — Code Aesthetics with Agentic Reward Feedback from October 2025, which released an AesCoder-4B model and an AesCode-358K dataset — and nothing in the AesCode-32B card cites it or states a relationship to it. If you go looking for background reading and land on that one instead, you are reading about a different model built by overlapping authors.
• Compare it on a public leaderboard — not yet. No third-party index appears to have evaluated it, which is unsurprising for a repository with two downloads.
Who should care, and who should wait
The audience for this is narrower than the headline table suggests and more specific than "anyone building with models." If your product turns prompts into decks, posters, reports or dashboards that somebody has to edit afterwards, you have already discovered the choice this model is aimed at: image generation gets you a beautiful rectangle you cannot change, and code generation gets you something editable that looks like it was assembled by a compiler. A 33B open-weights model that emits a full HTML document with real tables and ECharts specs, that stays within about a point of its reference-conditioned quality when you have no reference to give it, and that you can fine-tune on your own house style under Apache 2.0, is a genuinely useful thing to have exist. Nobody else's model does that exact job at this size with these terms.
Against that: everything you know about the quality comes from a table the vendor built, the rubric behind it was produced by the same pipeline that made the training data, and the one ceiling the card admits to — Style below 60 for every system tested — is precisely the dimension a design-sensitive product would care about most. A team with a compliance reason to keep generation in-house and a spare eight-GPU node has enough to start today. A team choosing a model for production next week has no independent number to choose on, and should not read the 85.89 against 81.28 as a settled result.
What would turn this into a story
Four things, none of which exists yet. An announcement — Microsoft has said nothing, and the paper the card references is explicitly under review, so a technical report with training details beyond the card's summary could appear at any point. An independent run — the reference-free robustness claim and the Boundary score, which the card says falls to a severe failure on 4.3% of samples against 34.7% for GPT-5.5, are both cheap to test and both worth testing. Serving support outside the vendor's recipe — a GGUF build or a mainstream runtime entry would change the hardware story more than any benchmark would. And a second data point on the size question: an 8B model scoring 82.94 Overall against the 32B's 85.89 on the same table is a two-point gap for a quarter of the parameters, which is the sort of thing that either gets reproduced or quietly stops being mentioned.
Until one of those lands, the accurate description of AesCode-32B is this: real weights under a permissive licence, a training stack detailed enough to be reproduced by a well-equipped lab, a benchmark table that is one company's measurement of its own model and of two competitors', and a repository history that will not let you call any single day the launch date. It is a more interesting artifact than its two-download counter suggests, and a less proven one than its 85.89 implies. Both halves of that sentence are the honest read on 11 October 2026.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
