การ์ดฮีโร่สำหรับ Ling-3.0-flash-base เป็นการ์ดชื่อเรื่องที่มีภาพประกอบสไตล์แฟลตของจุดตรวจสอบสามจุดที่ซ้อนกันแล้วรวมเป็นน้ำหนักฐานสุดท้ายหนึ่งเดียว โดยระบุป้ายกำกับว่า pre-trained, mid-trained และ merged ใต้หัวข้อหลัก Ling-3.0-flash-base และหัวข้อย่อย Ant Group's open checkpoints for continued pre-training
Engineering & Research

Ling-3.0-flash-base: การเปิดตัว checkpoint อย่างเงียบ ๆ ของ Ant Group ตอนนี้เป็นทางการแล้ว

ผู้เขียน

Elias Hawthorne

วันที่เผยแพร่

โมเดลล่าสุด · 20ดูโมเดลทั้งหมด
เบนช์มาร์ก: Artificial Analysis · อัปเดตทุกวัน
กลับไปยังโพสต์ทั้งหมด

ห้องแล็บ inclusionAI ของ Ant Group ได้อัปโหลด Ling-3.0-flash-base และเช็คพอยต์พี่น้องอีกห้าตัวไปยัง Hugging Face เมื่อวันที่ 11 สิงหาคม 2026 โดยไม่มีการประกาศใดๆ และตลอดหนึ่งสัปดาห์เต็ม ทั้งหกที่เก็บรหัส (repositories) มียอดดาวน์โหลดเป็นศูนย์ — ดูเหมือนไม่มีใครนอกแล็บสังเกตเห็น จากนั้นในวันที่ 19 สิงหาคม บัญชีทางการของแล็บได้เปิดตัวคอลเลกชันนี้ต่อสาธารณะ และการ์ดโมเดลทุกใบในคอลเลกชันได้รับการอัปเดตภายในไม่กี่นาที นี่คือบทความสรุปเท่าที่เรารู้ในตอนนี้เกี่ยวกับการเผยแพร่ครั้งนี้ สรุปสั้นๆ ก่อนเลย: Ling-3.0-flash-base ไม่ใช่โมเดลแชท และไม่ได้ตั้งใจให้เป็นเช่นนั้น มันคือเช็คพอยต์ฐาน (base checkpoint) ดิบก่อนการปรับ alignment ของ Ling-3.0-flash ซึ่งเป็นโมเดล Mixture-of-Experts ที่มีพารามิเตอร์ 124B และพารามิเตอร์แอ็กทีฟ 5.1B เผยแพร่ภายใต้ใบอนุญาต MIT สำหรับผู้ที่ต้องการฝึกมันต่อมากกว่าแค่เรียกใช้งาน

เส้นเวลา: เงียบหนึ่งสัปดาห์ แล้วจึงมีการประกาศ

ลำดับมีความสำคัญ เพราะมันบอกคุณว่าจะอ่านส่วนอื่น ๆ ในบทความนี้อย่างไร รีโพสิทอรีทั้งหกถูกสร้างขึ้นเมื่อวันที่ 11 สิงหาคม และน้ำหนักสำหรับ Ling-3.0-flash-baseซึ่งเป็น checkpoint ที่ถูกผสาน (merged checkpoint) และเป็นที่มาของชื่อบทความนี้ ถูกปล่อยพร้อมกับคอมมิต "initial model release" ในวันที่ 12 สิงหาคม ไม่มีบล็อกโพสต์ ไม่มีเธรดบนโซเชียล ไม่มีการประกาศใด ๆ มาพร้อมกับมัน หนึ่งสัปดาห์ต่อมา เมื่อถึงเวลาที่ทำการวิจัย รีโพสิทอรีทั้งหกยังคงแสดงยอดดาวน์โหลดเป็นศูนย์

ความเงียบนั้นสิ้นสุดลงเมื่อวันที่ 19 สิงหาคม บัญชีทางการของ Ant Ling ประกาศการรวบรวม checkpoint ต่อสาธารณะ และการ์ด Hugging Face ของทั้งหก repo ถูกอัปเดตภายในช่วงห่างกันไม่กี่นาที — การประทับเวลาตรงกันเกือบจะพอดี ดังนั้นการมองอย่างตรงไปตรงมาคือ: ตัวเวตถูกปล่อยออกมาอย่างเงียบ ๆ การประกาศมาช้าไปหนึ่งสัปดาห์ และทุกสิ่งที่รู้ได้ในตอนนี้เกี่ยวกับการเผยแพร่ครั้งนี้ มาจากการ์ดโมเดลเอง บวกกับคลังเก็บ training-cookbook สั้น ๆ ยังไม่มีสิ่งอื่นใดที่ได้รับการยืนยันโดยอิสระ

สิ่งที่ปล่อยออกมาจริง: checkpoint หกตัว, ขั้นตอนการฝึกสามขั้น

InclusionAI อธิบายการเปิดตัวครั้งนี้ว่า "กลุ่มของจุดตรวจสอบในระหว่างกระบวนการฝึก" สำหรับซีรีส์ Ling-3.0 ซึ่งเป็นตระกูลโมเดลพื้นฐานภาษาที่มีประสิทธิภาพสูงที่สุดจนถึงปัจจุบัน สำหรับขนาดที่เล็กที่สุดสองขนาด — Ling-3.0-tiny และ Ling-3.0-flash — ห้องปฏิบัติการเผยแพร่จุดตรวจสอบสามจุด โดยหนึ่งจุดต่อขั้นตอนการฝึก:

ผ่านการฝึกฝนล่วงหน้าLing-3.0-flash-base-30T และ Ling-3.0-tiny-base-30T: เสร็จสิ้นการฝึกฝนล่วงหน้าขนาดใหญ่แล้ว; ไม่มีการฝึกฝนขั้นกลาง, ไม่มีการรวมโมเดล, ไม่มีการฝึกฝนภายหลัง

ผ่านการฝึกขั้นกลางLing-3.0-flash-base-midtrain และ Ling-3.0-tiny-base-midtrain: เสร็จสิ้นการฝึกขั้นกลางแล้ว ไม่มีการรวมเช็คพอยต์ และไม่มีการฝึกภายหลัง

ผสานLing-3.0-flash-base และ Ling-3.0-tiny-base: เช็คพอยต์ฐานที่ผ่านการผสานด้วย WSM ซึ่งสร้างจากน้ำหนักที่ผ่านการฝึกระยะกลาง และแสดงถึงสถานะฐานสุดท้ายก่อนการฝึกขั้นหลัง

ทั้งหกเป็นเช็คพอยต์พื้นฐาน ยังไม่มีตัวใดผ่านขั้นตอน alignment — supervised fine-tuning, preference optimization, หรือ RL — ซึ่งเป็นขั้นตอนที่เปลี่ยนโมเดลพื้นฐานให้เป็นสิ่งที่สามารถนำไปให้ผู้ใช้ใช้งานได้จริง การ์ดโมเดลระบุไว้อย่างตรงไปตรงมาว่า: ไม่แนะนำสำหรับการแชทโดยตรงกับผู้ใช้ปลายทาง แอปพลิเคชันที่วิกฤตต่อความปลอดภัยโดยไม่มีการ alignment และการประเมินผลเพิ่มเติม หรือการใช้งานในโปรดักชันโดยไม่ผ่าน post-training และการตรวจสอบเฉพาะงาน

เบสเช็คพอยต์คืออะไร — และทำไมการพลิกผันของ WSM จึงสำคัญ

{{1}}หากคุณเคยใช้งานโมเดลผ่าน API เท่านั้น{{/1}} {{2}}คำว่า "base checkpoint" อาจฟังดูเหมือนโมเดลเวอร์ชันที่เล็กกว่าหรือเก่ากว่าที่คุณรู้จัก แต่จริงๆ แล้วไม่ใช่ทั้งสองอย่าง{{/2}} {{3}}base checkpoint คือไฟล์น้ำหนักของโครงข่ายประสาทเทียมที่ยังไม่ได้ผ่านการปรับเทียบalignment ซึ่งได้มาจากการฝึกก่อน (pre-training) ส่วนพฤติกรรมการทำตามคำสั่งและการเรียกใช้เครื่องมือที่คุณใช้งานจริงนั้นถูกเพิ่มทับลงไปบนไฟล์นี้ในภายหลัง{{/3}} {{4}}การเผยแพร่ base checkpoint คือวิธีที่ห้องปฏิบัติการมอบวัตถุดิบให้ชุมชนนักวิจัยเพื่อนำไปทำ continued pre-training, fine-tuning, distillation และ RL ต่อไปได้ด้วยตนเอง{{/4}}

สิ่งที่ทำให้รีลีสนี้มีความน่าสนใจมากกว่าการปล่อยน้ำหนักฐานทั่วไปก็คือสูตรการฝึกที่อยู่เบื้องหลังมัน InclusionAI กำลังใช้เทคนิคที่เรียกว่า วอร์มอัป-เสถียรและผสาน (WSM) และมันแทนที่การลดอัตราการเรียนรู้แบบเดิมตอนท้ายการฝึกด้วยการรวมเช็คพอยต์แบบถ่วงน้ำหนัก พูดง่ายๆ ก็คือ แทนที่จะปล่อยให้อัตราการเรียนรู้ค่อยๆ ลดลงเพื่อจบการรัน ห้องแล็บจะฝึกที่อัตราคงที่ บันทึกเช็คพอยต์ในช่วงสุดท้าย และรวมพวกมันเข้ากับน้ำหนักที่เรียนรู้มา เหตุผลที่ผู้พัฒนาระบุไว้ในบัตรโมเดล:

• เนื่องจากไม่มีระยะการสลายตัวที่จะ "ตรึง" ส่วนผสมข้อมูลสุดท้าย โมเดลฐานที่ได้จึงเหมาะสำหรับ การฝึกต่อเนื่องและการขยายข้อมูลแบบไดนามิก.

• เนื่องจากสามารถสำรวจสูตรการผสานได้แบบออฟไลน์ ทีมจึงสามารถเปรียบเทียบโปรไฟล์การเสื่อมสภาพต่างๆ โดยไม่ต้องรันการฝึกอบรมที่มีค่าใช้จ่ายสูงซ้ำ สำหรับแต่ละโปรไฟล์

ประเด็นที่สองนี้เองที่ทำให้การเผยแพร่ครั้งนี้เป็นสิ่งประดิษฐ์เพื่อการวิจัย มากกว่าจะเป็นเพียงของน่าสนใจ กล่าวคือ สาระสำคัญของ checkpoint เหล่านี้คือการนำไปฝึกต่อ และสูตรการฝึกถูกออกแบบมาให้รองรับการฝึกต่อได้

ภายใน Ling-3.0-flash-base มีอะไร

สถาปัตยกรรมเหมือนกับ Ling-3.0-flash หลังการฝึก ซึ่งน่าจะคุ้นเคยกันดีอยู่แล้วจากการกล่าวถึงโมเดลนั้นอย่างครอบคลุม จากโมเดลการ์ดและ config.json:

ขนาด — พารามิเตอร์ทั้งหมด 124B ตามที่ผู้ผลิตระบุ (เช็คพอยต์เต็มบนดิสก์ รวมถึงเอมเบ็ดดิงส์ นับได้ประมาณ 127B บน Hugging Face) โดยมี 5.1B ที่ถูกใช้งานต่อโทเคน.

ความเบาบาง — โมเดล MoE แบบเบาบาง 1/64: ผู้เชี่ยวชาญแบบกำหนดเส้นทาง 512 ตัว โดยเปิดใช้งานผู้เชี่ยวชาญแบบกำหนดเส้นทาง 8 ตัวและผู้เชี่ยวชาญร่วม 1 ตัวสำหรับทุกโทเค็น

Attention — a native hybrid-linear design: 35 Kimi Delta Attention (KDA) layers and 7 Gated MLA layers in a 5:1 repeating pattern across 42 transformer layers.

บริบท — คอนฟิกนี้มี position embedding ขนาด 256K โทเคน (262,144) ซึ่งเป็น native context เดียวกับรุ่นพี่น้องที่ผ่าน post-training แล้ว

น้ำหนัก — BF16, มาเป็น 28 ชาร์ด safetensors (~255 GB บนดิสก์) พร้อมกับแบบกำหนดเอง bailing_hybrid โค้ดโมเดลที่รวมอยู่ในที่เก็บ.

ใบอนุญาต — MIT โดยไม่มีข้อกำหนดเพิ่มเติมว่าด้วยการใช้งานที่ยอมรับได้

Single-model scoreboard for Ling-3.0-flash-base listing six figures: total params 124B (5.1B active), stage base pre-alignment WSM-merged, context 256K native, architecture hybrid-linear MoE 1/64 sparse, experts 512 routed with 8 plus 1 active, and independent scores none yet, with the footer noting all figures are from the vendor model card

One honest caveat about the size figure: the vendor's model card says "Total 124B," and that is the number that circulates everywhere — but the total safetensors parameter count on the repository is about 127.5B once the embedding layer is included. The 124B figure is the non-embedding total, and the two numbers are not in conflict; it is worth knowing which one you are looking at when you see "127B" in a sidebar or a downloader.

What it is for, and what it is not

The model card's recommended use cases are unambiguous:

• Continued pre-training

• Mid-training

• Supervised fine-tuning for domain adaptation

• Preference optimization and RL post-training

• Distillation research

• Long-context and MoE systems research

And its "not recommended as-is" list is just as explicit: direct end-user chat deployment, safety-critical applications without alignment, production use without post-training. In other words, this is raw material for people who train models, not a model people deploy. If you have no plans to fine-tune or pre-train anything, there is nothing for you to download here — and nothing wrong with that.

The other half of the story is the scale-seamlessly design. InclusionAI states that Ling-3.0-tiny-base and Ling-3.0-flash-base share the same training recipe, so the community can validate a continued-pretraining or fine-tuning strategy on the tiny checkpoint first — which fits in far less GPU memory — and then scale the same validated recipe to the larger flash checkpoint. That is a deliberate workflow choice, not an accident of release logistics, and it is the most useful property of the collection for anyone who actually wants to experiment.

Hardware reality, and the route most people actually want

There is no getting around the size of the flash base weights. BF16 at ~255 GB means a serious multi-GPU node (or a memory-optimised training setup) before you even start; there is no official FP8 or FP4 quantisation of the base checkpoint — the quantised variants that exist in the org's catalogue are for the post-trained model. The fine-tuning examples live in the inclusionAI/ling-cookbook GitHub repository, which is explicitly a work in progress. If continued pre-training is not your team's focus, the base checkpoints are not your entry point.

Most people reading about the Ling-3.0 family do not want a base checkpoint at all — they want the model that talks. That is the post-trained Ling-3.0-flash, available through the vendor's own API and several third-party platforms. When a fast, newly released open-weights model like that appears, the cheapest way to judge it against your current stack is to run it side by side with what you already use — and that is exactly what a router is for: one API, one key, and automatic failover so a provider hiccup does not end the trial. OrcaRouter connects to 200+ models this way and passes provider list prices straight through at 0% markup, so the price you see is the provider's own price, and a vendor cut lands on our side the same day.

Screenshot of the inclusionAI/Ling-3.0-flash-base model card on Hugging Face showing the six-checkpoint collection table for Ling-3.0-tiny and Ling-3.0-flash and the training-stage descriptions

What's not yet confirmed

Because this release is a day old in announcement terms, the unconfirmed list is long and should be read as a to-do list, not a criticism:

No independent benchmarks of the base checkpoint. The model card embeds a vendor benchmark chart comparing Ling-3.0-flash-base with pretrained peers across math, coding, reasoning, multilingual, and long-context suites — but it is the vendor's own evaluation, hosted on an internal image server, with no methodology section and no third-party reproduction.

No community reproductions yet. As of research time, all six repos showed zero downloads and essentially no engagement. The release is too fresh for anyone to have published an independent fine-tune or continued-pretraining run on top of these weights.

Tooling is early. The bailing_hybrid architecture ships with custom modeling code, so you are depending on trust_remote_code or on a framework that has added support. The fine-tuning cookbook is explicitly work in progress.

The announcement itself is thin. The public introduction came as an X thread and the model cards; there is no standalone launch document describing the training data, the compute, or the exact merge recipe beyond what the cards state.

By contrast, the post-trained sibling does have an independent track record: Ling-3.0-flash is tracked by Artificial Analysis at an Intelligence Index of 38 — a figure that comes from the leaderboard, not from the vendor. None of that exists yet for Ling-3.0-flash-base; a base checkpoint is not served anywhere, so it is not on any leaderboard. The two artifacts are at completely different stages of verification, and it helps not to blur them.

Screenshot of the Artificial Analysis model page for Ling-3.0-flash, the post-trained sibling of Ling-3.0-flash-base, showing its Intelligence Index of 38 and pricing of $0.075 per 1M input and $0.22 per 1M output tokens

ใครควรลงมือตอนนี้

Three groups should read this differently:

Researchers doing continued pre-training, domain SFT, or distillation — this is the release for you. The WSM-merged base checkpoint is specifically built to be trained on further, and the tiny-then-flash recipe gives you a cheap way to validate an approach before scaling it. Start with Ling-3.0-tiny-base, prove the recipe, then move to Ling-3.0-flash-base.

Engineers who want a working model — you want the post-trained Ling-3.0-flash, not this release. The base checkpoints will not follow instructions and are not served anywhere; do not point an API client at them.

Everyone else — the useful move is to watch. The interesting signals over the next few weeks are a first independent fine-tune built on these weights, a community reproduction of the vendor benchmark chart, and whether a third-party training stack adds first-class support for the bailing_hybrid architecture.

คำถามที่พบบ่อย

What is the difference between Ling-3.0-flash-base and Ling-3.0-flash?

Ling-3.0-flash is the post-trained model — aligned with supervised fine-tuning and RL, follows instructions, does tool-calling, and is served through APIs. Ling-3.0-flash-base is the same architecture at an earlier stage: it is the final merged base checkpoint before any post-training, and it will not behave like a chat model. The base exists so that researchers can do their own continued pre-training, fine-tuning, or distillation rather than starting from scratch.

Why would a lab release checkpoints that have not been post-trained?

Because post-training is exactly what most researchers want to do themselves, and a base checkpoint is the correct starting point for it. Releasing the pre-trained, mid-trained, and merged checkpoints separately means a team can pick the stage that matches their work: continued pre-training from the pre-trained weight, or fine-tuning and RL from the merged base — and, thanks to WSM, without the learning-rate-decay constraints that make a base model fragile when you train it further.

Can I run Ling-3.0-flash-base on consumer hardware?

Not practically. The BF16 weights are roughly 255 GB across 28 shards, with no official quantised variant of the base checkpoint. The smaller sibling, Ling-3.0-tiny-base, is the one intended for experimentation on more modest hardware — and it shares the flash checkpoint's training recipe, so validated experiments scale up.

Is Ling-3.0-flash-base free to use commercially?

Yes — it is released under the MIT license with no acceptable-use rider. That covers both research use and, in principle, commercial use of the weights and of any fine-tunes built from them. It is a base checkpoint, so "use" here means training on it; the license says nothing about a hosted API, because none exists for the base model.

บรรทัดล่าง

For most readers, Ling-3.0-flash-base is a pointer, not a product: it tells you where the Ling-3.0 family is headed and how the lab intends the community to build on it, but there is nothing here to call through an API. For the smaller group actually doing continued pre-training, fine-tuning, or distillation, it is a genuinely useful release — six checkpoints, MIT-licensed, deliberately designed to be trained on further, with a tiny-first path that makes experimentation affordable. What would move it from "useful release" to "proven recipe" is exactly what is missing today: an independent reproduction, by someone outside the lab, of what WSM checkpoints can do when somebody keeps training them.

© 2026 OrcaRouter

สำหรับผู้ให้บริการ

ให้บริการแพลตฟอร์มการอนุมานอยู่หรือไม่ นำโมเดลของคุณขึ้น OrcaRouter

providers@orcarouter.ai

เข้าร่วมคอมมูนิตี้ของเรา

Discordsupport@orcarouter.aiXGitHubYouTube