Blog
Engineering deep-dives, product updates and AI infrastructure notes from the OrcaRouter team.
Engineering & ResearchFeaturedLFM2.5-2.6B-Base: Liquid AI's Quietest Release Is the One Fine-Tuners Actually Wanted
Liquid AI shipped LFM2.5-2.6B-Base with 34T training tokens, 151 downloads and no published benchmarks. What the repo, config and license really say.
Guides & InsightsOpenAI Astra: Everything We Know About the Model That Isn't GPT-6
OpenAI named its next flagship Astra on Aug 1, 2026 — ten Lean proofs, no date, no price. What's confirmed, what's rumor, and when it could ship.
Guides & InsightsInkling-Small: The Quarter-Size Model That Beats Its Own Flagship, and Still Trails the Index
Inkling-Small beats the 975B Inkling on Thinking Machines' own benchmarks at a third of the cost, yet trails it on the independent index. The real numbers.
Engineering & ResearchMotif-Audio: The Audio Foundation Model Motif Technologies Shipped Without Telling Anyone
Motif Technologies shipped Motif-Audio with no announcement and a pre-release TODO still in its card. What the weights really do, and what they cannot.
Guides & InsightsMicrosoft Mage-VL: A Codec-Native 4B Video Model, Shipped Without an Announcement
Microsoft quietly shipped Mage-VL, a 4B codec-native VLM that reads compressed video, not frames. What's proven, what's unverified, who should care.
Engineering & ResearchIntern-S2-Mobius: The 35B Model That Separates Knowledge From Reasoning
Intern-S2-Mobius is a 35B open-weight model that splits knowledge from reasoning for up to 4.6x throughput. Benchmarks, architecture, and the catch.
Guides & InsightsGemini 3.5 Flash-Lite vs Qwen 3.8: Cheap Workhorse vs 2.4T Flagship
Qwen 3.8 is GA now at $2/$6, so this is a real tier fight: Flash-Lite is 6.7x cheaper on input, Qwen brings 1M context and multimodal.
Guides & InsightsGrok Voice Think Fast 2.0: xAI's Fastest, Priciest Voice Agent, Explained
xAI's Grok Voice Think Fast 2.0 answers in 0.70s and tops the agentic voice benchmark - but costs 60% more. Pricing math and the Aug 5 migration.
Engineering & ResearchGemini Robotics ER 2: Google DeepMind's Embodied-Reasoning Robot Brain, Explained
Google DeepMind's Gemini Robotics ER 2 explained: an embodied-reasoning VLM that plans, reasons about space, coordinates robots, and streams in real t
Guides & InsightsSeedance 2.5 vs Luma Ray 3.2: Reference-Heavy Production vs Cinematic Motion
Seedance 2.5 vs Luma Ray 3.2: reference-heavy long-form production vs Luma's cinematic motion and keyframe control. Length, control, and motion compar
