
Claude Opus 5.5 Demos Cost $3 to $26: How to Read the First Days of Builds
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 495 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 186 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1306 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 113 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 224 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
- metaMeta: Muse Spark 1.22026-08-0540Intelligence72Coding
Two numbers from the first week of Claude Opus 5.5 describe the same model and disagree by an order of magnitude. One builder reports a finished interactive lens lab — a page where you drag a focus ring and watch a glowing plane of focus sweep through a valley — at $25.66 in API spend and one hour 26 minutes of wall clock. A widely-shared round-up of launch demos prices a procedurally built Blender castle animation, complete with fireworks and a lake reflection, at $13.30 and 35 minutes. A thread that made the front page of r/ClaudeAI says $3.21. The vendor shipped Claude Opus 5.5 on September 22, 2026 at $4 per million input tokens and $20 per million output, and none of those three figures is wrong. They are measuring different things, and if you are budgeting against them you need to know which.
The gap is not noise. It is the difference between what an agent run produces and what it consumes, and the last four days have produced a remarkable amount of the latter with very little accounting for it. So here is the arithmetic nobody posted alongside the demos, then the part of the launch that is actually on a rate card and can be checked.
What $25.66 actually bought
The expensive end of the range is the most useful end, because the builder published both the artifact and the bill.
Ryan Sael's lens lab, titled "The Plane of Focus," is a single interactive page built from one prompt in one session. It puts a camera lens on screen, lets you drag the focus ring, and shows you the thing every photographer knows about but almost nobody has seen: the thin sheet in the world where a photo actually turns sharp. Turn the ring and the sheet moves, the glass elements slide forward or back inside the barrel to show why, and a console reports the distance, the aperture, and the sharp zone — 60 cm, f/2, 1.9 cm of depth, at the settings the page opens on. Stop the aperture down from f/2 to f/16 and the sharp zone widens as the cone of light narrows. An exploded view separates the glass. There is no image model anywhere in the pipeline, and no downloaded assets: every element is drawn and animated in code, which is why the whole demonstration of the circle of confusion is a live simulation rather than a diagram.
That is a genuinely good explainer, and the claim attached to it is unusually precise: one shot, 1 hour 26 minutes, $25.66 of API usage, with the page live at lens.lab.sael.net when I read it. The figure is self-reported and unauditable, but it is also the only demo cost in circulation that comes with a session length to divide by.
Run the arithmetic on the rate card. At Claude Opus 5.5's published $4 per million input and $20 per million output, $25.66 is between 1.28 million tokens and 6.4 million tokens depending on the mix — the upper end if the run were nothing but input, the lower end if it were nothing but output. An agentic build like this sits somewhere in the middle, and the reason it sits closer to the low end is the line item most people skip: cache reads on Claude Opus 5.5 cost $0.20 per million, a fortieth of the input rate, so a long session that keeps re-reading the same growing codebase pays almost nothing for that traffic. Ten million cached input tokens costs $2. The output side is where an hour and a half of a model writing code, thinking, and rewriting it accumulates: at $20 per million, a million output tokens is $20, and thinking tokens bill as output. Sael's number is consistent with a run that produced on the order of a million output tokens against a large, cheap, cached input side — which is exactly what an unattended 86-minute build looks like.
Two things follow. First, the sticker price cut that got the attention on launch day, from $5/$25 to $4/$20, is a 20% change per token and would not, on its own, have made this run cheap; the cache-read cut from $0.50 to $0.20 is the line that made a 90-minute loop affordable. Second, and less comfortable: nothing in the $25.66 report tells you whether the run went well. A well-executed one-shot and a session that thrashed for an hour and a half and then succeeded produce the same invoice.
The $13.30 figure does not survive its own arithmetic
The middle of the range is instructive for a different reason. A launch-demos round-up published September 23 reports a head-to-head in which Claude Opus 5.5 was asked to procedurally build a ten-second Blender animation from one prompt — castle, lake reflection, fireworks — against GPT-6 Astra, with no manual editing, and reports 35 minutes and 199,600 output tokens for about $13.30, against 28 minutes and about $14.50 for GPT-6 Astra. The token count is the interesting part, because it does not reconcile with the price.
199,600 output tokens at Claude Opus 5.5's $20 per million is $3.99, not $13.30. To reach $13.30 on output alone you would need 665,000 tokens. The missing money is in the other columns: input at $4 per million, cache writes at $5 per million for a five-minute breakpoint and $8 per million for an hour, and any retries. That is not a contradiction in the report — it is the report being honest about a token count that counts one thing and a bill that counts four, which is the single most common confusion in launch-week cost talk. If you take one number from a published demo, take the dollar figure, and treat any token count beside it as a partial description of the run.
At the cheap end, the accounting gets thinner still. The r/ClaudeAI thread that reports $3.21 of API usage for a finished build is typical of the genre: a title, a video, and no breakdown. It is plausibly a shorter session, a smaller artifact, or a lot of cached input. It is not something to plan a budget against, and I could not verify a dollar of it.
Why every one of these numbers is a lower bound
The failure mode here is not that builders dissemble. It is selection. The demos that circulate are the ones that worked, published by the people they worked for, and the four attempts that produced a broken page cost the same as the one that did not. The round-up that compiled the Blender comparison says this outright — that nobody posts the failures — and its own most detailed entry is a case in point: a watch-brand site that missed the brief and used an unreadable font on the first pass and needed a second to become the version that got shown.
"One shot" is doing quiet work in that vocabulary too. It describes a single session, not a single attempt, and a session is allowed to contain as much self-correction as the model can generate — which, on a 86-minute run, is a lot of it. The model rewriting its own code twice inside one context window is one shot and two attempts.
So the honest reading of the first four days is three tiers of evidence, and they should not be averaged:
• Self-reported cost with a session length — usable for order of magnitude, unverifiable. Sael's $25.66 over 1h26m is the strongest item in this tier.
• Self-reported cost with a token count — internally inconsistent more often than not, as the $13.30 figure shows.
• A dollar figure with nothing behind it — a signal that someone built something, and no more than that.
The part of the launch that is on a rate card
Strip the demo claims and the verifiable picture is narrow and clear. All figures below are Anthropic's published rates, read from the company's own model documentation and announcement on September 25, 2026:
• Input — $4 per million tokens. Output — $20 per million, with thinking tokens billed as output, and Claude Opus 5.5's adaptive thinking always on and impossible to switch off.
• Cache reads — $0.20 per million, a 0.05x multiplier. Cache writes — $5 per million at the five-minute breakpoint, $8 per million at one hour. Batch API — 50% off both directions, $2 and $10.
• Fast mode — $8 input and $40 output per million, up to 2.5x the speed, and Anthropic still labels it a research preview.
• Envelope — 1M-token context, 128K output tokens synchronously and 300K on the Batch API with a beta header, reliable knowledge cutoff of June 2026, default effort medium rather than the high that Claude Opus 5 defaulted to.
• Availability — the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry, and Claude Platform on AWS, with retirement committed no sooner than September 22, 2027.
The vendor's own performance claims, labeled as claims: Claude Opus 5.5 performs at the level of Claude Fable 5.1 on most work, costs about 40% less to run than Claude Opus 5, and generates output more than 30% faster. The 40% is a workload-average the vendor selected, not a rate-card number, and the two should not be quoted as if they were the same kind of fact.
The independent check is smaller than the vendor table and points the same direction with less margin. Artificial Analysis scores Claude Opus 5.5 at 57.6 on its Intelligence Index, alongside 61.4 on Humanity's Last Exam, 66.9 on SciCode and 84.7 on long-context recall in the figures its model page carries. On the same publication's coding-agent measurement, which reports cost per task rather than cost per token, the number moves the other way: at maximum effort the cost per task rises to about $13.04, against $10.79 for Claude Opus 5 in the same harness. A cheaper token that the model spends more of is not a cheaper task. That is the same lesson the demo bills are teaching, arriving from a third party.
What to do differently in week two
If the $25.66 number is the one that made you want to try this, the useful move is not to budget $25.66 — it is to make your own run produce a number you can explain.
Set effort explicitly rather than accepting the default, because it is the dial that moves tokens per task more than the model choice does, and the default on this model changed from Claude Opus 5. Cache the stable prefix deliberately: the minimum cacheable prompt on Claude Opus 5.5 is 512 tokens, down from 1,024, so short shared prefixes that never cached before now do. Give any long-running job headroom in max_tokens, since thinking is part of that budget and can no longer be turned off — Anthropic's own migration guidance suggests starting at 64K for work at the highest effort levels and tuning from there. And instrument the run at the end, not the start: read the top-level model field on every response, because on prompts its classifiers flag, Anthropic re-routes the request inside the same call, and the model string you sent does not describe what answered.
That last point is where a router stops being a convenience. A build that runs unattended for an hour and a half through a safeguard boundary can come back having silently finished some turns on a smaller model, and the only way to find out is to look. Claude Opus 5.5 is on OrcaRouter at Anthropic's own list price with 0% markup — provider rate passed straight through, so a vendor price change is live here the same day it is published — which is the least interesting thing routing does here. The interesting thing is that one API for 200+ models puts the whole experiment on a single key, so "run it on the frontier model, and fall back rather than fail" is a configuration line instead of a second contract and a rewrite. Automatic failover is what makes an unfamiliar model safe to put on a path you care about, and a refusal you can route around is a refusal that does not end your run at minute 70.
The number that will move next
Anthropic has said Claude Sonnet 5.5 and Claude Haiku 5.5 follow "in the coming weeks," and that will reprice every one of these demos without changing their shape: a cheaper per-token rate against a token count nobody can predict in advance. The same uncertainty that made the first week's numbers unusable will apply to the next generation's, which is the reason to build the habit now rather than after the price cut lands again.
For the moment the honest summary is short. Claude Opus 5.5 is real, generally available, and priced at $4 and $20 with the cheapest cached reads the line has ever carried. The demos it produced in its first four days are remarkable and their price tags are anecdotes. If you want a cost for your own workload, the only instrument that reads it is your own run, capped, with the invoice read afterwards.



Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
