Một thẻ tiêu đề được tạo tự động ghi "Boreal vs Kimi K3" cùng phụ đề "Cả hai đều không chạy trên phần cứng của bạn", các biểu tượng đường nét phẳng gồm một khung phát video, một chồng máy chủ dạng thanh và một ổ khóa, với logo OrcaRouter được ghép ở góc dưới bên phải.
Guides & Insights

Boreal vs Kimi K3: Không cái nào chạy trên phần cứng của bạn

Tác giả

Magnus Corvin

Ngày đăng

Mô hình mới nhất · 20Xem tất cả mô hình
Benchmark: Artificial Analysis · cập nhật hằng ngày
Quay lại tất cả bài viết

Ninety-six Safetensors shards. That is what downloading Kimi K3 gets you — a 2.8-trillion-parameter mixture-of-experts model with 104 billion parameters active per token, published under a modified MIT licence on 27 July 2026, eleven days after Moonshot AI released it. Moonshot's own guidance for serving it recommends supernode configurations of at least 64 accelerators. Boreal cannot be downloaded at all: Creatify's advertising video model, launched 15 September 2026, is closed and self-served, priced at one cent per second of finished 720p-class video. One of these hands you the weights and the hardware bill. The other keeps the weights and charges you by the second. Neither one is running on your hardware, and the reasons are opposites — which is exactly why the comparison is worth making.

What "open weights" actually ships you

The open-weights label does real work in this comparison, and it does less work than it sounds like.

What you genuinely get with Kimi K3 is durability. The weights exist as files. If Moonshot AI changed its pricing, deprioritised the model, or disappeared, the checkpoint would still be there — and so would the serving stacks that support it, with vLLM, SGLang and TokenSpeed all listed among the supported paths. That is not a marketing claim; it is a property of the artifact. For an organisation whose risk register includes vendor concentration, it is the single most valuable thing about the release.

What you do not get is a deployment you can run. A 2.8-trillion-parameter sparse model with 896 experts is not a quantise-and-go proposition. The official recommendation of 64-plus accelerators is not a conservative baseline that a determined team can undercut by much — it is a description of the class of machine the model was designed for. Teams that run Kimi K3 at anything resembling production scale are, in practice, renting it from someone who operates a supernode. Which means the practical availability question is unchanged by the weights being public: you are still buying inference from a third party, and the licence is the only thing that moved.

There is a licence detail worth reading closely, because it is the part most summaries skip. The modified MIT terms gate branding and attribution obligations above 100 million monthly active users or $20 million in monthly revenue, and require model-as-a-service businesses whose licensee-and-affiliate revenue exceeds $20 million over any twelve consecutive months to enter a separate agreement with Moonshot. Internal use and fine-tuning are broadly permitted. Resale is gated. That is a reasonable licence for a frontier model, and it is a materially different thing from MIT — for a startup planning to build a product on top of it, the threshold is a future negotiation, not a present right.

What Boreal ships you instead

Boreal makes the opposite trade, and it is more deliberate than it looks.

Creatify states that it owns the weights and serves them itself, and that the base model it post-trained from is an open-source one that generates audio and video jointly. Those two facts together are the entire product strategy. The open-source lineage tells you roughly what class of model is being called and that the audio is not bolted on afterwards. The ownership tells you that no one else can serve it, which is what makes a per-second price possible in the first place.

What you get is not durability — it is throughput. Creatify quotes rendering at roughly 1:1 with realtime: one second of finished video per second of wall clock. That is an inference-engineering result rather than a model-quality result, and it is the thing that would survive the weights being published tomorrow. If the checkpoint appeared on Hugging Face next week, the ability to render a five-second clip in about five seconds for a nickel would still belong to whoever built the serving stack — and the reference implementation, with 64-plus accelerators' worth of throughput already tuned, would still be Creatify's.

So the honest symmetry is this: Kimi K3 gives you a checkpoint you cannot afford to serve, and Boreal gives you a service you cannot take away from its vendor. Both are un-self-hostable. Only one of them is a hedge against vendor risk, and it is the one where the hedge costs more than the thing being hedged.

A generated two-column scoreboard for Boreal vs Kimi K3 sharing six dimension labels. Boreal reads output rendered video, weights closed and vendor-served, self-hosting not possible, meter $0.01 per second, context not published, independent score none yet. Kimi K3 reads output text only, weights open with 96 shards, self-hosting 64+ accelerators, meter $3 / $15 per 1M, context 1M at a flat rate, independent score 1679 Elo front-end. Footer reads "Boreal figures vendor-run and unreproduced; Kimi K3 per our catalogue and third-party evaluators."

So sánh thông số kỹ thuật

• Output — Boreal: rendered video, 720p-class, text-to-video and image-to-video. Kimi K3: text only.

• Weights — Boreal: proprietary, owned and served by Creatify. Kimi K3: published 27 July 2026, 96 Safetensors shards, modified MIT with revenue-gated resale terms.

• Self-hosting — Boreal: not possible. Kimi K3: possible in principle, with Moonshot recommending at least 64 accelerators.

• Price — Boreal: $0.01 per second of finished video. Kimi K3: $3.00 per 1M input / $15.00 per 1M output, with cached reads at $0.30 per 1M and a flat rate across the full context window.

• Context — Boreal: not published. Kimi K3: 1,048,576 tokens, with no long-context surcharge.

• Latency — Boreal: quoted at 1:1 with realtime. Kimi K3: p50 time to first token of 10.00 seconds on our board — the ceiling our instrumentation reports, so read it as "slow" rather than as a precise measurement.

• Evidence — Boreal: vendor-run only. Kimi K3: vendor rows including 93.5 on GPQA Diamond and 88.3 on Terminal-Bench 2.1, a first place in the Frontend Code Arena at 1679 Elo, and independent findings that include a 36% failure rate on problems with hidden invariants.

The trace you cannot switch off

Kimi K3 has a billing property that is unusual even among reasoning models, and it is the most concrete cost difference in this comparison.

Thinking mode cannot be disabled, and the default reasoning effort is set to maximum. Every request produces a full reasoning trace, and every token of that trace bills at $15.00 per million output. For a task where the answer is a single word, you are still paying for a deliberation. Cache reads at $0.30 per million soften the input side considerably, and the flat pricing across the full 1M context is genuinely friendlier than the tiered structures its competitors use — but the output side has no cheap mode.

Boreal's cost has no equivalent lever to get wrong. A five-second clip is five cents whether the model worked hard or trivially, because the artifact is the unit and the artifact is fixed. That is not obviously better — it is simply unaffected by anything the model does internally.

Kimi K3 is in our catalogue, and it is the busiest model on our board by a wide margin: 3,150.6 million tokens over the last seven days, more than seventeen times Grok 4.6's traffic and thirteen times GPT-5.6 Sol's. That volume is what an open-weight frontier model looks like in production — many callers, price-sensitive, willing to route around whoever is serving it worst. It sits behind the same key as the rest of the catalogue at provider list rate with 0% markup, which matters more for an open-weights model than a closed one: when the checkpoint is public, the price is set by whichever provider is serving it, so the price is the competitive variable and it moves.

Boreal is not in our catalogue. OrcaRouter does not serve video generation models, and the text layer of an advertising pipeline — the briefs, the hook variants, the compliance pass — is where we sit relative to it.

A screenshot of Creatify's own launch post for Boreal, dated September 15 2026 and headed "Introducing Boreal: frontier-quality AI video at a cent a second", showing the Boreal by Creatify Labs logo card and the opening paragraph stating that Boreal generates text-to-video and image-to-video at one cent per second, about $0.05 for a five-second clip, and renders in realtime.

Which one, and what to watch

If your constraint is vendor concentration and you have access to serious compute, Kimi K3 is the only one of these two that offers a real answer, and the 36% failure rate on hidden-invariant problems is the number to weigh against that. It is a frontier-class model with a genuine escape hatch attached, and the escape hatch is expensive enough that most teams will never use it.

If your constraint is cost per finished advertising asset, Boreal is the cheaper and far more predictable option, and its constraint is quality on one specific format: a single-person talking clip scored 60% in the vendor's own 40-case test, against 83% and 85% on product and creator scenes. Test that format before you commit a campaign to it.

One caution on the evidence, because it applies to both. Kimi K3's published intelligence figures have been restated at least once, and older numbers in the 57-to-60 range should not be quoted alongside current ones from the same evaluator — a reminder that a benchmark score is a dated observation, not a property of the model. And every quality figure attached to Boreal comes from Creatify's own launch post. The first independent evaluation of Boreal, whenever it arrives, will be worth more than everything currently published about it.

A screenshot of the OrcaRouter model page for kimi/kimi-k3, showing Kimi K3 by MoonshotAI dated 2026-07-15 marked Featured, a 1M-token context, text + image input with text output, a p50 time to first token of 10.00 seconds, input $3.00 and output $15.00 per 1M tokens, and 3150.6M tokens of traffic in the last 7 days.