A generated hero card titled 'Claude Haiku 5.5 vs Nemotron 3.5 Lightning 30B A3B Base' with the kicker 'THE CHEAP ONE CANNOT ANSWER A QUESTION' and chips reading $0.10 vs $0.06 per million input tokens, AA Index 43 against 13, and answer-ready against base checkpoint.
Guides & Insights

Claude Haiku 5.5 vs Nemotron 3.5 Lightning 30B A3B Base: The Cheap One Can't Answer a Question

Author

Elias Hawthorne

Date Published

Latest models · 20View all models →
Benchmarks: Artificial Analysis · updated daily
Back to all posts

On the two numbers a procurement spreadsheet cares about most, the cheaper and faster model wins. Nemotron 3.5 Lightning 30B A3B Base runs at 298.9 output tokens per second, the fastest of 142 models tracked independently, and costs $0.06 per million input tokens against Claude Haiku 5.5's $0.10. It is also, as the word "Base" in its name is warning you, not an assistant. It is a pre-trained checkpoint — no supervised fine-tuning, no reinforcement learning, no chat template worth the name — published on the OpenMDW-1.1 licence on August 11, 2026 so that other people can build on it. Claude Haiku 5.5, released October 7, 2026, is a finished product you can send a question to. Comparing them on price is like comparing the cost of flour to the cost of bread, and the interesting part of this matchup is working out which one your workload actually needs.

Nobody in the launch coverage says that part clearly, because "cheaper and faster than Claude Haiku 5.5" is a better headline than "cheaper, faster, and not usable without a fine-tuning run." Both statements are true. Only one of them is useful.

What each model actually is

Claude Haiku 5.5 — Anthropic, ID claude-haiku-5-5, released October 7, 2026. Text and image in, text out. 1M-token context, 128,000-token maximum output (300,000 behind a beta header on the Batch API), June 2026 training cutoff, retirement no sooner than October 7, 2027. Adjustable effort across Low, Medium, High, Xhigh and Max, defaulting to Medium, with adaptive thinking on by default. Metered API, cache reads at $0.01 per million, Batch at half rate. Available through Anthropic's API and, from there, wherever Anthropic's models are resold.

Nemotron 3.5 Lightning 30B A3B Base — NVIDIA, announced August 11, 2026. A 30-billion-parameter hybrid mixture-of-experts checkpoint with roughly 3 billion active parameters per token, built on a Mamba-2 plus MoE plus attention stack, distilled from Nemotron 3 Ultra, and shipping with MTP, DFlash and DSpark speculative decoding. Context runs to 1M tokens, though about 256,000 is the practical ceiling on a single H100. It is text-only. The weights are open under OpenMDW-1.1. And it is a base checkpoint: raw pre-trained weights with no instruction tuning, which means it completes text rather than following instructions.

A generated two-column scoreboard titled 'Claude Haiku 5.5 vs Nemotron 3.5 Lightning 30B A3B Base — the scoreboard'. Claude Haiku 5.5 reads output speed 243.4 tokens per second, AA Index 43, proprietary weights, instruction-tuned yes, $0.10 per million input and $0.21 per task; the Nemotron column reads output speed 298.9 tokens per second, AA Index 13, weights open under OpenMDW-1.1, instruction-tuned no — base checkpoint, $0.06 per million input and $0.09 per task. A footer notes the Artificial Analysis figures are independently run and the NVIDIA figures are vendor-reported and unreproduced.

• The dimensions, one line each

• Input price — Claude Haiku 5.5 $0.10 per million to 100K tokens, then $0.50; Nemotron 3.5 Lightning 30B A3B Base $0.06 per million.

• Output price — $0.50 per million to 100K tokens then $2.50, against $0.20 per million.

• Cost per finished task — $0.21 per independent index task for Claude Haiku 5.5, against $0.09 for Nemotron — the cheaper model is cheaper per task, not just per token.

• Output speed — 243.4 tokens per second for Claude Haiku 5.5, ninth of 182; 298.9 tokens per second for Nemotron, first of 142.

• Time to first token — about 0.3 seconds for Claude Haiku 5.5, against 0.54 seconds for Nemotron.

• Independent score — Artificial Analysis Intelligence Index 43 for Claude Haiku 5.5, second of 182, with the Max configuration at 43.40; Index 13 for Nemotron, twenty-ninth of 142.

• Context window — 1,000,000 tokens each, with roughly 256,000 practical for Nemotron on a single H100.

• Modalities — text and image in for Claude Haiku 5.5; text only for Nemotron.

• Weights — proprietary against open under OpenMDW-1.1, a 30B total and 3B active hybrid MoE.

• Instruction tuning — present on Claude Haiku 5.5, absent on the Base checkpoint.

The "Base" problem, stated plainly

A base checkpoint predicts the next token. Ask it a question and it will often continue the text in the style of a document that might contain such a question, rather than answering. Prompt it with "Summarise this contract" and a base model may produce a plausible continuation of a legal-training document rather than a summary of yours. NVIDIA publishes a separate BF16 checkpoint and separate instruction-tuned siblings precisely because the base weights are raw material.

On NVIDIA's own model card — vendor-reported, not independently reproduced — the base checkpoint scores 78.59 on MMLU, 67.94 on MMLU-Pro, 91.28 on GSM8K, 77.44 on HumanEval and 69.62 on RULER at 1M tokens. The Lightning card for the tuned sibling reports MMLU-Pro 81.94, GPQA Diamond 75.44, SWE-Bench Verified 51.56 and 86% on PinchBench, with roughly 30% faster inference than the model it replaced. Even the tuned numbers are NVIDIA's own, and the base numbers are the ones that apply to the checkpoint named in this article's title. Against them, Claude Haiku 5.5's independently measured 43.40 — second of 182 on Artificial Analysis' index, a different index measuring different things — is not directly comparable, and the honest reading is that a 30B base checkpoint and a finished small model are being scored by different instruments for different purposes.

What can be said without qualification is the shape of the gap. A base checkpoint that needs a fine-tuning pipeline, a serving stack and an evaluation harness before it answers anything is not a substitute for a hosted model at $0.10 per million. It is a starting point for someone who wants to own the model.

Where Nemotron genuinely wins, and it is not a consolation prize

Speed first. 298.9 output tokens per second puts it at the top of a 142-model board — 23% faster than Claude Haiku 5.5's 243.4. For a latency-sensitive pipeline, that is a real, measurable difference, and it is the reason the "Lightning" name is on the box.

Open weights second, and this is the bigger one. Under OpenMDW-1.1 you can download the checkpoint, fine-tune it on your own data, and serve it yourself. Claude Haiku 5.5 cannot be fine-tuned at all and cannot be self-hosted at any price. For a team with a proprietary task, an unusual input distribution, or a compliance requirement that traffic never leaves their own infrastructure, that difference is decisive regardless of what the benchmark pages say. A 30B total with 3B active is also small enough to be economically servable — the architecture is designed for exactly that.

Price third: $0.06 input and $0.20 output per million, with a 17% cache discount and $0.09 per index task, is roughly half of Claude Haiku 5.5's $0.21. Both models are cheap; Nemotron is cheaper.

Where Claude Haiku 5.5 wins, which is every dimension that requires an answer

Instruction following, because it has been trained for it. Independent capability, at Index 43 against 13 — a thirty-point gap on a board that scored both, which is the largest such gap in this set of comparisons. Image input, which Nemotron does not accept at all. Maximum output at 128,000 tokens against an unpublished figure, with a 300,000-token beta path on Batch. And the effort dial, which lets you buy back cost by lowering reasoning effort rather than by rebuilding the model.

The verbosity caveat cuts both ways here. On Artificial Analysis' suite, Claude Haiku 5.5 emits about 440 million output tokens against a median of roughly 100 million — it thinks a great deal. Nemotron emits about 120 million, which the same source describes as somewhat verbose for its size. Haiku 5.5's cost per task is nonetheless $0.21 against Nemotron's $0.09, so even the chatty model is not the expensive one. What that tells you is that per-token prices mislead in both directions, and the per-task number is the one to budget against.

Self-hosting versus one endpoint

Nemotron 3.5 Lightning 30B A3B Base is distributed as open weights and served by several inference providers; NVIDIA also publishes the tuned siblings. Claude Haiku 5.5 is available through Anthropic's API and, from there, wherever Anthropic's models are resold. Neither is on OrcaRouter's catalogue, and we will not pretend otherwise — what we route from this neighbourhood is anthropic/claude-haiku-4.5, anthropic/claude-sonnet-5.5, anthropic/claude-opus-5.5 and anthropic/claude-fable-5.1, the tier above this article's subject.

The two paths this comparison describes have genuinely different plumbing, and it is worth naming them. The self-hosted path means a GPU, a serving framework, a quantisation decision and an ops rotation. The hosted path means a key and a model string. OrcaRouter exists for the second path: one API for 200-plus models, one key and one bill, provider list price passed through at 0% markup so a vendor cut is live the same day, with automatic failover underneath. It is not a route to a 30B base checkpoint, and buying a DGX-class box is not a route to a hosted small model.

A capture of Anthropic's Claude Platform models documentation showing the current Claude lineup — Claude Fable 5.1, Claude Opus 5.5, Claude Sonnet 5.5 and Claude Haiku 5.5 — with each model's context window, maximum output and per-million input and output pricing.

Who should pick which

If your task is well-defined, your training data is large enough to matter and your traffic cannot leave your own infrastructure, Nemotron 3.5 Lightning 30B A3B Base is the more interesting object: fastest of 142 on output speed, half the price per token, and yours to modify. Budget for the fine-tuning run and the serving cost before you count the savings, because the $0.06 input rate assumes somebody has already done the work that turns a checkpoint into an assistant.

If you want to send a prompt and get an answer this afternoon, Claude Haiku 5.5 is the one that does it — 43 on the independent index against 13, image input, a 128,000-token output ceiling and a price that is genuinely small. The comparison only looks close on a price sheet. It stops looking close the moment the task is to answer a question rather than to continue a document.

A capture of the Artificial Analysis page for Nemotron 3.5 Lightning showing its Intelligence Index of 13 and rank of 29 of 142, the $0.06 per-million input and $0.20 output pricing, the $0.09 cost per index task, the 298.9 tokens-per-second output speed ranked first of 142, the 0.54-second time to first token, the 1M-token context window, text-only input and open weights.