
Claude Opus 5.5 Multimodal: It Renders Video From Code, and Cannot Watch One
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 1009 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAINEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 195 tok/s
- OrcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1189 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 22 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 108 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 220 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- DeepSeekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- xAISpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
Ask Claude Opus 5.5 for a video and you can end up with one — a program of HTML, CSS or canvas calls that paints every frame, which you then run. Hand it a video and ask what happens in it and you get nothing at all, because there is no video input. That asymmetry is the whole of this model's modality surface, and it is identical on all eleven Anthropic models on our board, so it is worth stating flatly: Claude Opus 5.5 reads text, images and files, writes text, and stops there. The board around it is far less uniform. MiniMax H3 generates video outright, OrcaDub 1.0 takes a video in and returns a dubbed one, and Qwen3.8-Max will watch a clip for two dollars per million input tokens — half what Claude Opus 5.5 charges for the same million, and with a capability it does not have.
This is a reading of the modality line as OrcaRouter records it on 2026-09-29, across all 204 models in the catalogue. The numbers below are catalogue declarations, not our benchmarks — a modality field is what the model card says, and model cards move.
What the catalogue actually records
Of the 204 models, 182 declare an input surface. The other 22 leave it blank, which is a statement about the record rather than about the model — eight Kling video models, four OrcaRouter meta-models, and a scattering of free-tier and preview endpoints among them.

Across those 182:
• 131 read images
• 77 read files — and every one of those 77 also reads images, so file input is a strengthening of image input on this board, never a separate sixth modality
• 49 read video
• 19 hear audio
• 50 take text and nothing else
Claude Opus 5.5 is on the image-and-file line and not on the video line. Our own model page says it in one sentence, under the heading "Context, modalities and thinking": multimodal input — text, images and files — with text output, a 1M-token context window and up to 128K output tokens. That is the complete input surface. There is no fifth entry hiding in a submenu.
Fifty is a larger number than most people expect. More than a quarter of the models that declare anything at all take text only, which means a pipeline built on the assumption that it can attach a screenshot and have it understood will break on a quarter of this board.
Why "video with code" is not a video modality
The demos that made the rounds are real, and so is the effect: Claude Opus 5.5 is unusually good at writing a program that draws motion — a self-contained renderer that produces frames when you run it. That is a genuine capability and a useful one. What it is not is an output modality. The catalogue records exactly one output for this model, and it is text. The video is manufactured downstream, by the code the model wrote, on your machine, with your renderer.
The practical consequence cuts both ways, and the second half is the half people miss.
On the way out, you own the render step. Code-based video is inspectable, diffable, re-runnable at a different resolution, and cheap in the way a text response is cheap — a few dollars of tokens rather than a per-second billing meter. You also own every failure: a missing font, a bad frame budget, a renderer that disagrees with the code it was handed.
On the way in, there is nothing to hand it. If your source material is a screen recording, a product clip, a two-hour webinar, or a competitor's launch film, Claude Opus 5.5 cannot look at it. You get to send the transcript, the slides as still frames, or a description someone else wrote. That is the constraint that actually decides architectures, and it is invisible in a demo reel.
Two camps: 60 models take the document, 29 take the clip
Group the 182 models by their exact input set and the board separates into two groups that barely overlap.
Camp A — documents and images, no video: 60 models. Median input price $2.00 per million tokens. This is where Claude Opus 5.5 sits, alongside all eleven Anthropic models, three of the five Grok rows, and the reasoning end of the OpenAI catalogue — every GPT-4.1 and GPT-6 row, and in the GPT-4o and GPT-5 families everything except the voice, codex, search and chat-latest variants. Camp A also holds the ceiling of the price table: openai/gpt-5.5-pro and openai/gpt-5.4-pro at $30 per million input tokens. That is the highest input price anywhere on the board, and it is a price they share with two OpenAI GPT-4 rows and two text-to-speech endpoints, none of which reads a document. Not one of the eight can watch a video.
Camp B — text, images and video, no files: 29 models. Median input price $0.25 per million tokens. Twenty-two of the twenty-nine carry a Qwen prefix; the rest are MiniMax M3, GLM-5.3-Flash, Kimi K2.6, Kimi K2.7-Code, Gemma 4 26B and two Obsidian-tuned Qwen builds. Its ceiling is Qwen3.8-Max at $2.00 per million — which is to say the most expensive video-reading model in this camp costs exactly half of Claude Opus 5.5.
The median gap between the camps is 8×. The membership gap is starker still. Of the 182 models that declare an input surface, 60 take documents and not video, 32 take video and not documents, 73 take neither, and only 17 take both. The overlap is small enough that the two groups read as nearly disjoint products, and the 17 in the middle are almost entirely Google's upper tier — fifteen of the seventeen — plus Meta's two Muse Spark builds.

There is a plausible reason for the shape. Document understanding and video understanding are different engineering problems with different cost curves, and the vendors who solved video first are, mostly, the ones selling cheap tokens by the hundred million. The vendors who solved documents first are the ones selling reasoning at a premium. Nobody has yet been forced to do both at the same price, so the market has split into two products and priced them accordingly.
The nineteen that take everything
Nineteen models accept all four input types — text, image, video and audio. Sixteen are Google, two are Meta, one is MiniMax. Google's own range inside that camp runs from $0.10 per million for gemini-2.5-flash-lite to $4.00 for gemini-pro-latest, so "reads everything" is not a price tier; it is a decision Google made at every price point. Meta's contribution is the pair of Muse Spark builds — Muse Spark 1.1 and Muse Spark 1.2, identical in the catalogue at $1.25 per million in and $4.25 out, with a 1M-token context.
This is the camp that makes the operative point of the article: a video-aware pipeline does not require a flagship. It requires picking from a list of nineteen, ten of which cost less than a dollar per million input tokens.
Five models on this board emit something other than text
Reading video and making it are different tricks, and the second is rarer. Of the 204 models, 139 declare an output surface; five of those name something other than text.
Three are image generators: Nano Banana (Gemini 2.5 Flash Image) at $0.30 in and $30.00 out per million tokens, Nano Banana Pro (Gemini 3 Pro Image Preview) at $0.24 per request, and Nano Banana 2 (Gemini 3.1 Flash Image Preview) at $0.151 per request. Two emit video: MiniMax H3, billed per second of generated output — $0.08 per second at 768P and $0.13 at 2K, for clips of four to fifteen seconds with native stereo audio — and OrcaDub 1.0, billed at $0.60 per minute of source video, covering 28 languages and cloning a separate voice per speaker.
Two framings to avoid here. First, the remaining 65 rows leave the output field blank; blank means the catalogue does not record an output surface, and it is not evidence that a model emits only text. Second, reading video and making video are separate axes with almost no overlap on this board — OrcaDub 1.0 takes video in and gives video back, which makes it the one model here that plausibly sits at both ends of a pipeline, while MiniMax H3 takes four input types to produce a single output type that is not text.
What to route where
Four rules fall out of the census, and they are cheap to apply.
• If your input is a clip, do not reach for a frontier flagship first. The camp that reads video without needing documents is 29 models, twenty-two of them Qwen, and its median is a quarter per million. Send the flagship the frames and the transcript instead, and let it do the reasoning.
• If your input is a PDF, a contract or a long document, the video camp is the wrong camp. Only 17 models declare both file and video input, and every model that declares file input also declares image input, so the two travel together. Claude Opus 5.5 at $4.00 per million buys the document surface plus a 1M-token window, which is the trade you are making.
• If you need media out, you are choosing from five models, not from the board. Three make still images and two make video, and those two bill per second and per minute rather than per token — so a cost estimate built from input prices will be wrong by construction.
• If you need video out and your source is a script, code generation remains the cheapest route on this board. Claude Opus 5.5 writes the renderer for the price of text; MiniMax H3 renders it for eight cents a second. Chaining the two costs less than either alternative and leaves you a file you can edit.
FAQ
Can Claude Opus 5.5 watch a video?
No. Its declared input modalities are text, image and file. Video is not among them, and neither is audio. A screen recording or a product clip has to be converted — sampled frames, a transcript, or both — before it can be sent.
Does Claude Opus 5.5 generate video directly?
Not as a modality. It returns text, and that text is frequently a self-contained program that renders motion when you run it. The rendering is yours; the model never emits a pixel. If you want a model on this board that returns video itself, MiniMax H3 and OrcaDub 1.0 are the two.
Is "multimodal" a price tier?
No, and the board shows why in both directions. Google ships full four-input multimodality from $0.10 per million input tokens to $4.00, while the joint-highest input price on the board — openai/gpt-5.5-pro at $30 per million — reads documents and images but not video. What you pay for is the reasoning tier, not the modality count.
Why do so few models read documents if so many read images?
Because files and images travel together on this board: all 77 file-reading models also read images, so file input is an extension of image input rather than an independent capability. The number to remember is that 60 of the 182 models declaring an input surface take documents and images and no video at all — a third of the board, and where the flagship tier sits.
How should I price a video pipeline?
By the unit the model bills in. Video readers are priced per token, like any chat model. Video generators on this board are priced per second of output (MiniMax H3: $0.08 at 768P, $0.13 at 2K) or per minute of source (OrcaDub 1.0: $0.60). Mixing the two units into one estimate is the most common way a video budget comes out wrong.
The line to remember
The interesting thing about Claude Opus 5.5 is not that it is multimodal. Most of the board is. It is that the modality line falls in an unusual place — documents and still images in, text out, no video in either direction — while the demos that made it famous are video. Those two facts are not in conflict, and reading them as a conflict is what makes the code-video story confusing. The model writes a renderer; the renderer makes the film. What the model can be handed is a script, a document and a screenshot.

Everything above is a snapshot of the OrcaRouter catalogue on 2026-09-29 and the modality declarations in it. Vendor specs move, and a model card's modality list is the first thing to re-check before you build on a number. If a claim here matters to a decision, open the model page — the fields are on it.
