
Solar Mini 4 vs LFM2.5 2.6B Base: Three Times the Index, and a Cost Comparison That Does Not Exist
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-29$2.00 / $10.00 per 1M tokens
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-28$2.00 / $10.00 per 1M tokens · 159 tok/s
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 220 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 118 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1148 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 48 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 102 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 217 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
The honest way to compare Solar Mini 4 and LFM2.5 2.6B Base is to admit up front that they are not in the same market. Solar Mini 4 is Upstage's 35-billion-parameter sparse mixture-of-experts, released 22 September 2026, activating about 3 billion parameters per token, proprietary, and billed at $0.10 per million input tokens and $0.40 per million output. LFM2.5 2.6B Base is Liquid AI's 2.69-billion-parameter pretrained checkpoint, published to Hugging Face in early August 2026 under the LFM Open License v1.0, small enough to sit in memory on a phone. One is an endpoint. The other is a weight file, and if you are reading this because you want to run a model rather than call one, the comparison is already over.
What makes the pairing worth an article anyway is the one dimension where they genuinely overlap: both have been positioned against the same independent yardstick, and the shape of that comparison is not what the parameter counts predict. Artificial Analysis scores Solar Mini 4 at 24 on the Intelligence Index v4.3.2 and rates the LFM2.5 2.6B generation at 8.4 — but the second figure is an estimate, and it is not measuring the checkpoint in this article's title. Both caveats matter, and they are the first thing to sort out.
The spec sheet, side by side
• Parameters — Solar Mini 4 holds 35B total with 3B active; LFM2.5 2.6B Base is dense at 2.69B, with nothing held back. Thirteen times the weights on the Upstage side for roughly the same per-token compute budget.
• Architecture — Solar Mini 4 is a sparse mixture-of-experts. LFM2.5 2.6B Base is 30 layers, 22 of them double-gated short-convolution blocks and 8 grouped-query attention, the hybrid design Liquid built for CPU and edge inference.
• Training — Upstage does not publish a token budget. Liquid's card documents 34 trillion tokens and a 128,000-token vocabulary across 16 languages, Korean and Japanese among them.
• Licence — Solar Mini 4 is proprietary. LFM2.5 2.6B Base ships under the LFM Open License v1.0 — permissive for commercial use with conditions, and fine-tunable.
• Context — Upstage documents 512K tokens, Artificial Analysis lists 1.05M for the same model, so treat the ceiling as unsettled. LFM2.5 2.6B Base runs 131,072.
• Modality — both are text-in, text-out. Neither takes images or audio.
• Price — Solar Mini 4 at $0.10 / $0.01 cached / $0.40 per million tokens, currently 50% off through 22 October 2026. LFM2.5 2.6B Base has no rate card at all; the host Artificial Analysis tracks prices the LFM2.5 2.6B generation at $0.00.
What "8.4, estimated" does and does not measure

Artificial Analysis's entry for the LFM2.5 2.6B generation — the post-trained model, not the pretrained Base — carries an index of 8.4 flagged as an estimate, with most of its component evaluations marked the same way. That is a position on a leaderboard, not a specification to build against, and it should never be quoted as though someone ran the Base checkpoint through the suite. Nobody has. Liquid's own model card publishes no benchmark table for LFM2.5 2.6B Base at all: it is a pretrained text-completion checkpoint, and the card says plainly that it is recommended only for heavy fine-tuning.
The scored sibling is a different artifact. Liquid's benchmark table for post-trained LFM2.5 2.6B shows IFStruct at 85.5 and Multi-IF at 80.1, AIME25 at 51.9, LiveCodeBench v6 at 59.4, BFCLv4 at 56.9 and τ³-Banking at 5.7 — all vendor-run, all on the instruction-tuned build, and the τ³ figure is the one that reads like a warning.
Solar Mini 4's profile is a different shape entirely. It scores 83.3% on AA-LCR v1.1 for long-context reasoning, 47.6% on SciCode, and −10.8 on AA-Omniscience with 18.4% accuracy and a 64.2% non-hallucination rate. It also scores 1% on Terminal-Bench 4.0 and 22.3% on AutomationBench-AA. Both families are weak at agentic work; Upstage's model is weak at agentic work while being expensive, which is a worse place to be than being weak at agentic work while being free.
Where Solar Mini 4 earns its price is scale of reasoning. At 88,300 output tokens per index task — 71,600 of them reasoning — it spends compute in a way a 2.6B dense model cannot imitate by being handed a longer prompt. A model with 2.69B parameters can be told to think step by step; it cannot be told to have 35B parameters. That is the actual product.
Three times the index, and no comparable cost figure
Artificial Analysis puts Solar Mini 4 at $0.36 per completed index task against $0.006 for Granite 4.2 3B, and the LFM2.5 2.6B generation is priced at zero, so no cost-per-task figure exists that would make this a fair ratio. The framing "Solar Mini 4 costs vastly more than LFM2.5 2.6B Base" is therefore true and slightly beside the point. The relevant question is whether the Korean-language, long-document, structured-extraction work Solar Mini 4 is built for can be done by a 2.69B pretrained checkpoint at all. Usually it cannot, and then the comparison becomes Solar Mini 4 versus a frontier generalist rather than Solar Mini 4 versus a phone model.
The case where LFM2.5 2.6B Base wins is not the case where it is cheaper. It is the case where the task is narrow enough that fine-tuning closes the gap. Structured field extraction from a known document format, classification against a fixed taxonomy, an on-device assistant that must work with no network — those are tasks where a 2.69B checkpoint plus your own training data beats a 35B generalist plus a rate card, and costs a GPU instance instead of a token meter. Liquid ships the base checkpoint with continued-pretraining, LoRA SFT, DPO and GRPO recipes through Unsloth and TRL precisely for that motion.
Which one to call
Reach for Solar Mini 4 when the input is long, the output has to be reasoned over, and the language mix includes Korean or Japanese — and when seven minutes of decode per hard task is tolerable. Reach for LFM2.5 2.6B Base when the weights need to be yours, the device has no network, and the task is narrow enough to fine-tune. The interesting deployments use both, because they are not alternatives: a small local model does the routing, redaction and extraction, and a hosted model is called only for the fraction of requests that actually needed 35B parameters.
Liquid's model is not on OrcaRouter, and neither is Upstage's, so neither route is something we can sell you. What we can do is the part that turns this comparison from a decision into a configuration: more than 200 models behind one key at 0% markup over provider list price, with automatic failover across providers and a routing DSL that lets you compose several models into a single call. When the right answer is "the cheap one most of the time and the good one when it matters," that is a routing problem, and it is the reason the routing DSL exists.

The number neither vendor will put on a slide
Latency. Solar Mini 4 burns 88,300 output tokens per Intelligence Index task at a measured 204 tokens per second, so it averaged 430 seconds — a little over seven minutes — on each hard task. The LFM2.5 2.6B generation, on the host Artificial Analysis tested, ran at 45.7 tokens per second with a 2.8-second wait for the first chunk, but that is a datacenter host talking to a model that would ordinarily run on your desk. Liquid's own measurements put the same architecture at 220 tokens per second on an M5 Max, 113 on a Ryzen AI Max+ 395, and about 30 tokens per second on a phone — which is a usable agent loop on a device with no network.
If your workload is interactive, the comparison is not close and no benchmark score will change it. If your workload is batch and the thing waiting is a cron job, seven minutes is fine and the reasoning depth is worth it.

Two sourcing notes to close. Every Artificial Analysis figure here is that organisation's independent measurement on a shared suite; the LFM2.5 2.6B index is explicitly an estimate and covers the post-trained model rather than the pretrained Base checkpoint named above. Every Liquid AI figure is the vendor's own run on the vendor's own benchmarks, unreproduced — and the Base checkpoint itself has no published scores at all, which is a fact about pretrained models rather than a criticism of this one.
