
AesCode-8B vs North Mini Code: Same Licence, Same Download Button, Opposite Problems
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 87 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 115 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 47 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 777 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 61 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 452 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
The most useful number in this matchup is 150,000. That is the download count Cohere attached to North Mini Code on 2026-08-14, two months after the model launched on 2026-06-09, and it is the figure that makes the comparison with AesCode-8B legible. AesCode-8B, Microsoft's fine-tune of Qwen3-VL-8B-Instruct, shows two downloads on its Hugging Face repository. Both are Apache 2.0, both are ungated, both are downloadable this afternoon, and one of them has been in production hands for a quarter while the other has not left the lab. That gap is not a quality verdict — nobody has measured AesCode-8B independently, so there is nothing to be smug about on either side — but it is the honest starting point, because it tells you which one has a community you can ask questions of.
The two arrived very differently and it shapes what you can verify. North Mini Code is a Cohere and Cohere Labs open-weights release with a model card, a blog post behind it, a paper trail of harness choices, and a vendored API that served it at $0.00 per million tokens during the launch period. AesCode-8B has no announcement at all: Microsoft created the repository on 2026-09-29, committed the weights under the message "Release AesCode-8B" at 03:35 UTC on 2026-10-07, and pushed the training code to GitHub on 2026-10-08. No Microsoft Research post, no release note, no product page, and a citation line reading "Under review" with the placeholder year 2027. The 8B model emits slide decks, posters and dashboards as HTML and CSS; North Mini Code writes code inside an agent harness. They are not competitors so much as neighbours in the catalogue, which is its own kind of finding.
What each one was trained to emit
The clearest way to see that these two do not do the same job is to look at what comes out of them.
• Output — AesCode-8B: a complete HTML and CSS document, one rendered page per request. North Mini Code: text and tool calls, which in practice means patches, shell commands and agent trajectories.
• Size — AesCode-8B: 8.8B bf16 dense, in four safetensors shards totalling 17,543,339,408 bytes. North Mini Code: 30B total with 3B active parameters, a sparse mixture-of-experts with 128 experts and 8 active per token.
• Context — AesCode-8B: 24,576 tokens in the reported configuration, with training-time prompts and responses each capped at 8,192. North Mini Code: 256K in, 64K maximum generation.
• Inputs — AesCode-8B: text plus an optional reference image. North Mini Code: text only, with the harness supplying everything else.
• Trained on — AesCode-8B: 3,000 cold-start demonstrations and then GDPO reinforcement learning over 7,408 prompts for 400 steps, scored by six deterministic verifiers plus a visual rubric after rendering each candidate in a sandboxed browser. North Mini Code: agent harnesses — SWE-Agent, mini-SWE-Agent, OpenCode and Terminus 2 — so it works inside whichever one you already run rather than pinning you to one.
• Licence — Apache 2.0 both, weights included, no field-of-use carve-out, no contributor tier. On this axis they are identical and it decides nothing.
• Evidence — AesCode-8B: Microsoft's own 300-sample infographic rubric, unreproduced. North Mini Code: a vendor figure plus an independent measurement.
The evidence asymmetry, which is narrower than it looks
East side of this comparison has an outsider in it, and that is worth more than the benchmark gap. Artificial Analysis has run North Mini Code itself and scores it 33.4 on its Coding Index and 27.6 on its Intelligence Index — competitive with or above similarly sized open models. Cohere reports 61.0% pass@1 on SWE-Bench Verified and up to 2.8× the output throughput of Mistral's Devstral Small 2, both vendor figures. The independent run is the one that found the catch: North Mini Code emits roughly three times the output tokens of its size-class median to finish the same benchmark suite, around 75-77 million output tokens against a 25-38 million median. At the launch-period price of zero that costs nothing in cash. It costs latency, and it previews the bill once the promotional rate ends.
AesCode-8B has no equivalent outside measurement. Microsoft reports 82.94 Overall on its own 300-sample infographic evaluation, three generations per prompt with no selection, ahead of reference-conditioned GPT-5.5 at 81.28 and Claude Opus 4.8 at 80.39 and 31.1 Visual points clear of its own backbone. It also reports that reference-conditioned rewards lifted prompt-only Visual from 25.17 to 69.71, and that withholding the reference image at inference costs the model only 1.00 Visual point against 19.55 for the backbone. Those are unusually specific claims and the harness is open-sourced, which is more than most vendor tables offer. None of it has been reproduced by anyone. There is no arena listing, no leaderboard placement, no throughput measurement and no third-party pass over the rubric.
Note what that means for the comparison specifically: there is no shared number. One model is scored on whether software-engineering tasks pass and how many tokens they cost; the other is scored on whether an infographic's text stays legible and its layout stays inside the canvas. Putting 33.4 next to 82.94 compares a coding index to an infographic overall score.
Where each one gets deployed, and what that costs

North Mini Code's deployment story is the reason it passed 150,000 downloads. A 3B-active MoE quantises into roughly 20-24 GB of memory, Cohere's team demonstrated it on a Mac Studio through MLX, and the 3B active footprint is what makes local inference economical. It runs on a single H100 at full precision. Cohere has been serving it at $0.00 per million input and output tokens during the launch period, and offers a managed per-instance deployment route for production scale. The count itself is Cohere's own figure, aggregated across every distribution channel with no platform-by-platform breakdown, and the public Hugging Face monthly number is an order of magnitude smaller, which is consistent with the aggregate spanning the API, hosted gateways, package managers and local pulls. Read it as momentum rather than telemetry.
AesCode-8B's deployment story is a command line. The card gives it directly — vllm serve microsoft/AesCode-8B --limit-mm-per-prompt image=2 --max-model-len 24576, with transformers 4.57 or newer for the non-vLLM path. 17.5 GB of bf16 weights sit alongside a KV cache for 24,576 tokens and up to two images, so a single 24 GB card is technically enough and tight in practice; 40-48 GB is the realistic floor. Add a rendering stack if you want to score outputs the way the authors did, because every quality claim about this model was made by rendering its HTML in a browser with external requests blocked. There is no hosted endpoint that we can find, and it is not in our catalogue.
That last point is the practical one for anyone comparing these two on availability. North Mini Code lives behind the vendor's own API and several third-party platforms as well as the weights themselves. AesCode-8B lives only on Hugging Face and GitHub. If your constraint is "I need to call this from a service tomorrow," only one of these two answers yes.
The composition question neither model solves
Here is the trap in both halves of this comparison, and it is the same trap. Adopting either one means maintaining a self-hosted serving path for a specialist. AesCode-8B needs a multimodal stack and a renderer to be evaluated properly. North Mini Code needs a slot for a 3B-active MoE plus an agent harness around it, and its verbosity means the trajectory costs more than the token count implies. A team that wants both capabilities — a page generator and a repo agent — is now running two inference deployments, and for everything that is neither, a third.
That is the case where one endpoint in front of many models does real work rather than being a convenience. Keep the specialist on your own hardware where the economics and the data handling justify it, and route the routine traffic to whatever hosted model is best for it this week, all behind one key with provider list price passed through at 0% markup and automatic failover if a provider wobbles. The Qwen3-family backbone is a good illustration of how thin the line gets: Microsoft built AesCode-8B on Qwen3-VL-8B-Instruct, and vision-language members of that family are callable through OrcaRouter — Qwen3-VL-8B-Instruct at $0.18 per million input tokens and $0.70 output over a 131,072-token context, and Qwen3-VL 235B A22B at $0.40 and $1.60. Neither of those is AesCode-8B; the fine-tune is Microsoft's and no provider serves it. But if what you wanted from it was a model that reads a screenshot and writes code, and you did not specifically need the aesthetic-design fine-tune, the callable option costs a fraction of the GPU and takes an afternoon to try rather than a procurement cycle.

How to decide, given that neither is general-purpose
Take the licence off the table first, because it is a tie and it is not going to decide anything.
Ask what you want the output to be. If the answer is a rendered, editable page — a deck, a report, a dashboard — North Mini Code cannot help you and AesCode-8B is the only one of the two pointed at the problem. Go in with clear eyes: an unannounced research checkpoint, two downloads, a 24K context validated only on single infographic pages, an English-only language tag, and no independent evaluation anywhere. The Style ceiling in Microsoft's own table is 53.21 on a dimension defined as needing no further visual revision before delivery, which is the vendor telling you a human still edits the output.
If the answer is a change to a repository — a migration, a fix, a multi-step terminal job — North Mini Code is the one with harness training, the 256K window, the 3B-active deployment profile, a functioning vendor API, and an outside measurement of both its quality and its verbosity. Budget for the token count rather than the headline rate, because the independent finding is that it writes about three times as much as its peers to get to the same place.
And if your quarter contains both a design artifact and a repo migration, do not treat this as a two-model decision. Run the specialist where it earns its GPU, route the rest, and stop maintaining serving paths you did not need.

The difference that outlasts the benchmarks
One final thing separates these two, and it is not technical. North Mini Code has a vendor that published a card, a post, a Space to try it in and a methodology you can argue with, plus 150,000 downloads' worth of other people's experience. AesCode-8B has a repository, a README, and a citation that says under review.
That changes what the open questions even are. With North Mini Code the live question is whether its verbosity makes it uneconomic at scale once the promotional rate ends — and you can answer that by running it, because plenty of other people already have. With AesCode-8B the live question is what Microsoft intends to do with it: research artifact, first release of a product line, or something quietly superseded in six weeks. Nothing in the repository answers that, and no amount of reading the model card will. Watch for three signals — an official announcement, an independent reproduction of the 82.94, or a provider picking the checkpoint up — and treat the absence of all three as the current state of the evidence rather than a reason the model is uninteresting. The reason it is interesting is the 1.00-point reference drop, and that number is still sitting there waiting for someone other than Microsoft to check it.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
