
Claude Sonnet 5.5 vs Kimi K3: Neither Model Will Take Your Temperature Setting
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 984 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 109 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Here is a comparison where the two models agree on something that will break your code either way. Claude Sonnet 5.5, released September 28, 2026, rejects any request that sets temperature, top_p or top_k to a non-default value with a 400 error. Kimi K3, generally available from Moonshot AI since mid-July 2026, accepts no sampling parameters at all — no temperature, no top_p, no seed — and exposes reasoning depth only through a reasoning_effort control. Two different labs, two different architectures, one shared conclusion: the sampling knobs that most integrations still set out of habit are gone on both.
That is not the comparison anyone writes about. The comparison everyone writes about is price: Kimi K3 at $3 per million input tokens and $15 output against Claude Sonnet 5.5 at $2 and $10, which is a straightforward win for Anthropic on the rate card. But Kimi K3 is an open-weight 2.8-trillion-parameter mixture-of-experts model you can download and run inside your own boundary, and Claude Sonnet 5.5 is a closed model available on four clouds. Between a cheaper closed model and a more expensive downloadable one, the decision is not made on price per token at all.
What each one is, in one paragraph each
Claude Sonnet 5.5 is the second model in Anthropic's 5.5 generation, following Claude Opus 5.5 onto the Claude API five days after that launch. It ships as claude-sonnet-5-5 across the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry, keeps a 1M-token context window and a 128K synchronous output ceiling, and holds Claude Sonnet 5's pricing exactly: $2 in, $10 out, $0.20 per million for cache reads. Adaptive thinking is on by default at high effort. It is available with zero data retention, and Artificial Analysis gives it an Intelligence Index of 56 on its v4.3.2 revision — third of the 216 models it measures.

Kimi K3 is Moonshot AI's flagship: a 2.8-trillion-parameter mixture-of-experts model activating roughly 104 billion parameters per token, with a 1M-token context window and native visual understanding. It is built for long-horizon coding and multi-step knowledge work, and Moonshot positions it for programming-agent harnesses by name. It speaks the OpenAI API format, exposes reasoning_effort as its only reasoning control, and is open-weight under the Kimi K3 License — which is not the same thing as open-source, and the difference is the part to read before you build on it.
The rate card, and the number under it
Line the two up on the dimensions that decide a bill rather than a headline:
• Input price — Claude Sonnet 5.5 $2 per million vs Kimi K3 $3 per million
• Output price — Claude Sonnet 5.5 $10 per million vs Kimi K3 $15 per million
• Cache reads — $0.20 per million vs $0.30 per million
• Context window — 1,000,000 tokens vs 1,048,576 tokens
• Weights — proprietary, four hosted platforms vs open-weight under the Kimi K3 License
• Intelligence Index — Claude Sonnet 5.5 56 vs Kimi K3 44 on the same v4.3.2 revision
• Cost per Intelligence Index task — Claude Sonnet 5.5 $7.60 vs Kimi K3 $2.00
• Verbosity on the Index — Claude Sonnet 5.5 410M output tokens vs Kimi K3 160M
That last pair is the reversal worth sitting with. Anthropic's model is 50% cheaper on input and a third cheaper on output, and still costs 3.8 times as much to complete the same evaluation, because it spends 2.6 times the tokens getting there. On the composite it also scores twelve points higher, so this is not a case of a wasteful model buying nothing. It is a case of two models whose cost behaviour cannot be compared by reading two rate cards, which is why the index-task figure exists.
Neither of those token counts will be your token count. Moonshot's own documentation notes that K3's reasoning depth is set by reasoning_effort rather than by sampling, and the same is effectively true of Claude Sonnet 5.5's effort parameter — so on both models the largest single lever on your bill is an effort setting, not a price negotiation.

The 400 errors, in detail

The shared refusal of sampling parameters is the most practical overlap here, because it is the failure that shows up the morning after a model swap.
On Claude Sonnet 5.5 the rules are explicit and unforgiving. A non-default temperature, top_p or top_k returns a 400. Sending thinking: {"type": "disabled"} returns a 400 pointing at between_tools, the new lowest thinking setting, which is accepted at low, medium and high effort but rejected at xhigh or max. Forced tool use — tool_choice set to "any" or to a named tool — returns a 400 with the message "tool_choice: type \"tool\" and \"any\" are not supported for this model", and the fix is auto plus strict tool use or structured outputs. Still more: thinking blocks are now bound to the model and the account that produced them, so a conversation can be moved from Claude Sonnet 5 onto Sonnet 5.5 with its reasoning intact but cannot carry that reasoning out to a different model family, and an edited prefix can turn a replay into a 400.
On Kimi K3 the surface is smaller but the consequences are similar. There are no sampling parameters to send, so any client that hard-codes one is already sending a parameter the model does not read. Reasoning depth is a top-level reasoning_effort field instead.
Both changes point the same way: the deterministic-knob approach to prompt control is being replaced by an effort dial on the model side. Any pipeline that tuned itself with temperature 0.2 and a fixed seed has to be re-tuned by effort level on both of these models, and the re-tuning is the migration.
What open weights actually buy you here
The reason to consider Kimi K3 against a cheaper and higher-scoring closed model is not capability and not price. It is the boundary.
Running K3 inside your own infrastructure means prompts, documents and outputs never leave a network you control, which is a different compliance conversation from a zero-data-retention agreement with a vendor. It also means the model cannot be deprecated out from under you — Claude Sonnet 5.5 comes with an Anthropic retirement commitment not sooner than September 28, 2027, and a downloaded checkpoint has no such date. And it means the marginal cost curve flattens: after hardware, tokens are free, which is the only structure in which a 2.8-trillion-parameter model ever becomes the cheap option.
The licence is what you are actually signing. Kimi K3 is open-weight, not open-source, and Moonshot does not use the latter term. The Kimi K3 License is broad for most uses, but a model-as-a-service business above a rolling revenue threshold needs a separate commercial agreement, and products past large user or revenue thresholds must display "Kimi K3" attribution. Calling it MIT-licensed is a K2-era inaccuracy that still circulates. If you plan to build on the weights rather than call the API, this is a legal review rather than a footnote.
And the weights are large enough that self-hosting is a capital decision. A 2.8T-parameter MoE at roughly 104B active parameters is not a single-server deployment, and the effort required to stand it up is the real cost of the option — one that dwarfs the $5-per-million difference in output rates.
Where OrcaRouter fits
Most teams evaluating these two do not want to choose between a managed API and a hardware procurement. They want to route.
OrcaRouter is one OpenAI-compatible endpoint over 200-plus models with provider list price passed through and no markup, so a vendor rate change on either side lands the same day rather than at the next invoice cycle. Kimi K3 has been in the catalogue since July at Moonshot's own $3 / $15, with a measured p50 first-token time around eight seconds and roughly 371.9 million tokens moving through it in the last week. Claude Sonnet 5 — the model most Claude API traffic still runs on, at a measured 150 output tokens per second — has been routable since June 30 at Anthropic's $2 / $10.
Claude Sonnet 5.5 is not in our catalogue yet. Where that is true, say so and route around it: the vendor's own API is the way to it, and the way to test it without committing is a percentage of traffic behind automatic failover, with Claude Sonnet 5 or Kimi K3 as the fallback. If your reason for looking at K3 is the boundary rather than the rate, the weights are downloadable today and the API is a stopgap — which is a decision about your infrastructure, not about a model comparison.
The verdict, split by what you are actually buying
Choose Claude Sonnet 5.5 when the work is long-context and output-heavy, when twelve points of composite index is the thing being purchased, and when a managed API with zero data retention clears your review. Budget for the token habit — 410M tokens on the Index against K3's 160M is a real multiplier — and budget a day for the tool-use and thinking-block changes.
Choose Kimi K3 when the boundary is the requirement, when you want a checkpoint that no vendor can retire, or when your workload is agentic coding at high volume where $2.00 per index task against $7.60 compounds. Accept that it is a 2.8T-parameter model to host, that its licence has attribution and revenue clauses, and that it scores below the Anthropic tier on the composite.
What neither choice escapes is the first paragraph: the sampling parameters are gone on both, and the effort dial is the control that replaced them. Whichever of these two you pick, that is the change your code needs.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
