A title card for North Micro Vision Instruct reading 'North Micro Vision Instruct' with the subtitle 'Cohere's 2.4B Apache-2.0 document VLM' and a chip noting it shipped August 12, 2026, above three line icons for document scanning, chart reading and bounding-box grounding.
Guides & Insights

North Micro Vision Instruct: Cohere's Quiet 2.4B Document VLM, Explained

Author

Rowan Sterling

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

CohereLabs/North-Micro-Vision-Instruct — "North Micro Vision" for short — is a 2.4-billion-parameter vision-language model that shipped the quiet way: the weights appeared on Hugging Face on August 10, 2026, Cohere's announcement followed on August 12, and a day later there is still no hosted API, no price list, and no third-party leaderboard entry. North Micro Vision is built for one job — reading documents at native resolution, the way a person squints at a scanned A4 page rather than the way a caption model glances at a thumbnail — and its model card is unusually direct about what it is not: no tool calling, no reasoning loop, system prompts actively discouraged. This is a what-we-know-so-far piece, so every figure below is labeled: what the repository itself establishes, what Cohere claims but nobody has reproduced, and what the one third-party look published so far adds.

What Cohere actually shipped

Everything in this section is read straight from the repository: config.json, the README, and the model card. These are Cohere's primary sources, and for most of what is known about the model they are the only sources.

• Size — 2.4B parameters total: a ~400M vision encoder plus a 2B "North Micro LLM" language model, in bfloat16, about 4.97 GB of weights

• License — Apache 2.0, commercially permissive, weights and tokenizer open

• Vision encoder — custom-trained for native resolution, initialized from SigLIP 2 SO400M

• Language model — the in-house 2B North Micro LLM following the Command A+ architecture: three sliding-window attention layers interleaved with one global attention layer, grouped-query attention (16 query heads over 8 KV heads), tied embeddings, 262,144-token vocabulary

• Projector — "DeepStack": patch embeddings from several vision-encoder layers are injected into early language layers rather than prepended once, which is how small text and table geometry survive

• Resolution — native-resolution input with aspect ratio preserved, up to about 1654 × 2339 pixels, which is one A4 page at roughly 200 DPI

• Context — 128K text context on the language backbone; 8K validated multimodal context (details below)

• Languages — 11 supported: English, Chinese, German, French, Spanish, Italian, Portuguese, Hindi, Japanese, Korean, Arabic

The "Instruct" suffix matters. This is the instruction-tuned checkpoint — a chat and visual-QA model — and as of writing it is the only checkpoint. The model tree already lists eleven community quantizations and three fine-tunes, but no separate base variant has been published under a different name.

Why the architecture is worth a paragraph

Most 2-3B VLMs downscale your image to a fixed grid and lose the fine print. North Micro Vision's bet is the opposite: keep the resolution, spend the tokens. The vision encoder uses 2D RoPE plus learned 1D positional embeddings, and the DeepStack projector feeds patch embeddings into multiple language layers rather than a single prepend, which is how a densely packed table or a footnote survives. Positional encoding is interleaved mRoPE, so the visual token count scales with image resolution instead of being capped by a preset.

That choice is the whole product — and it has a cost. RohitAI's launch analysis estimates a native A4 page at maximum resolution works out to roughly 3,800 merged visual positions. Read that number next to the context caveat below: 3,800 visual tokens plus a prompt can fill a large share of the validated multimodal window before you have asked your question.

The benchmark card, read honestly

A single-column scoreboard for North Micro Vision Instruct listing six rows: Size 2.4B (2B LM + 400M vision), License Apache 2.0, DocVQA 0.921, ChartQA 0.808, Multimodal context 8K validated, Tool calling none, with a footer noting the benchmarks are vendor-reported and not independently reproduced.

Every figure in this section is vendor-reported. None has been reproduced by an independent lab, and the model does not appear on any public leaderboard yet. On paper the record is genuinely strong where Cohere points it:

• DocVQA — 0.921 (document visual question answering)

• ChartQA — 0.808 (chart reading)

• OCRBench — 0.792 (OCR in the wild)

• RefCOCO grounding — 0.732 average P@1 (pointing at objects)

• MMMU — 0.329 (college-level multimodal knowledge — the weak corner, as expected at this size)

• IFEval — 0.749 (instruction following)

Cohere's launch framing claims North Micro Vision outperforms Gemma 4 E2B and Ministral 3 3B "on a range of visual understanding benchmarks," with the strength concentrated in document understanding and visual QA. That is Cohere's claim about Cohere's model — treat it as such. One wrinkle: Cohere's comparison against Liquid's family used the 1.6B checkpoint, not the 3B one, so there is no valid North-vs-LFM2.5-VL-3B comparison in the card even though the two models will be competing for the same edge-document budget.

The one third-party look published so far — a single blogger's run, not a lab suite — adds the picture the card leaves out. It reports North Micro Vision beats Ministral 3 3B Instruct on five of seven document, chart and OCR rows, which matches the vendor framing. But it also reports Qwen3.5-2B beats North Micro Vision on six of those seven rows and on every disclosed general-VQA, reasoning, counting, hallucination and text-knowledge row; North Micro Vision's only win over Qwen3.5-2B in that set is AI2D. And English OCRBench v2 comes in at 0.367 against the headline OCRBench 0.792 — a reminder that OCR scores depend heavily on which split and which difficulty you run. One run is not a verdict, but it is the first evidence that North Micro Vision's real competitor at this size is the Qwen 3.5 family, not the models Cohere chose to benchmark against.

One context window, two numbers

The card lists 128K of text context and 8K of validated multimodal context, and the gap between the two is doing real work. The 128K number belongs to the language backbone alone; image-plus-text prompts beyond 8K tokens were never trained or validated, so anything past that line is extrapolation. With a dense A4 page eating roughly 3,800 visual positions, you can reach that 8K window with one document and a short question. The practical consequence: plan your pipeline around cropping pages into regions — the table, the signature block, a single invoice — rather than feeding whole pages and expecting a long back-and-forth over them.

What the model card tells you not to do

North Micro Vision is a perception layer, not an agent, and Cohere is refreshingly explicit about it. The card says the model is not a reasoning model, has limited math and code ability, does not support tool calling or agentic workflows, and should not replace a larger general assistant. It goes a step further than most cards: it recommends against system prompts, because the model was not trained with them. There are no calibrated confidence scores either. The design this implies is a split: North Micro Vision reads and extracts, a schema owns structure, rules own the invariants, a stronger model handles the ambiguous calls, and a human approves anything consequential. Treat the text on the page as data, never as policy — a scanned instruction buried in a document is exactly the kind of visual prompt injection this split defends against.

A four-step document pipeline diagram reading Scan, North Micro Vision — perception, Schema + rules, and Stronger model — decisions, connected by arrows, with a footer reading 'Keep perception and decisions separate.'

How to run it today

The Hugging Face model card for CohereLabs/North-Micro-Vision-Instruct, showing the Image-Text-to-Text pipeline badge, Apache 2.0 license, 11 languages, the model tree, and the opening description of a 2.4B-parameter open-weight vision-language model with native-resolution image support, noting it is not deployed by any inference provider.

The quickstart is honest about its own friction: the model card requires Transformers 5.16.0, which was not the latest stable PyPI release on launch day, so early users install from source. Public vLLM support is listed as "coming soon"; Cohere evaluated with an internal vLLM build. Day-one ecosystem support is otherwise good: NVIDIA NeMo AutoModel recipes for edge and GPU deployment, Axolotl QLoRA and full fine-tuning recipes, and community MLX quantizations for Apple Silicon — the 4-bit build is about 2.17 GB. One deployment trap worth naming: set max_pixels explicitly, because sequence_len does not cap image tokens, and without it you can blow the validated multimodal window without meaning to.

North Micro Vision is not on OrcaRouter, and by its own card it is on no inference provider — this is a self-host story, full stop. Which is exactly why the routing layer matters in the next sentence: the moment any provider stands up an endpoint for it, a router is what lets you put a days-old Apache-2.0 model next to the ones you already call in production — one API, no new contract, provider list price passed through at 0% markup, and automatic failover so an unproven model can take real traffic while a proven one catches the calls it fumbles. Until that endpoint exists, the comparison you can actually run today is the one between this checkpoint and the hosted document models you already pay for, side by side on your own pages.

Who should move, and who should wait

Move if you have a bounded document-extraction job — invoices, bank statements, contracts, structured forms — and data-residency or audit requirements that make a hosted API unattractive. Apache 2.0, cheap QLoRA domain adaptation, and a native-resolution encoder are a genuinely unusual combination for that workload, and the card's own limitations are clear enough that you will not be surprised. Wait if you need a hosted endpoint, per-token pricing, or agentic behavior — none of those exist yet. And wait before treating any benchmark above as settled: every headline figure is Cohere's, the one external run paints a more contested picture at the top of the small-model class, and the model has been out for less than a week. The thing to watch is the serving story — the day stock Transformers and a public vLLM release both carry it, North Micro Vision stops being a repo you have to wrestle and becomes a model you can simply point at your documents.