
DeepSeek V4.1 Flash in OpenCode: Config, Cost, and the Effort Dial Nobody Mentions
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 151 tok/s
- OpenAINEWOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAINEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 126 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1202 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 52 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 251 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 230 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
DeepSeek V4.1 Flash has been sitting in OpenCode Go since 10 September, and the multiplier OpenCode attached to it has a published end date: 20 September. That is four days from now. The interesting part is not the deadline — it is that the setting which decides whether this model is pleasant to code with is missing from every setup guide on the first page of results, and in at least two harnesses it is silently dropped. This is a playbook for pointing a coding agent at DeepSeek V4.1 Flash: the exact model ids, the three integrations verified today, the community-reported reasoning-effort control, and the overthinking failure mode with the mitigations that actually work.
What you are actually pointing your agent at
DeepSeek V4.1 Flash shipped on 2026-09-10 — the release note on DeepSeek's own API documentation is dated that day, and it is the smallest model in the vendor's new architecture family. DeepSeek reports a 552B-parameter mixture-of-experts backbone built on what it calls a Causal Encoder–Decoder design, activating roughly 8B parameters per token during prefill and 16B during decode. It takes text and images as input and returns text, serves a context window of up to one million tokens, and will generate up to 384,000 tokens in a single response — the context and output ceilings both come from OrcaRouter's own model page for the model, which lists 1,048,576 and 384,000 respectively.
The vendor benchmarks, from the chart on DeepSeek's release page and therefore vendor-reported and unaudited: 30.0 on Terminal-Bench 3.0, 74.2 on DeepSWE v1.1, 88.1 on CyberGym, 54.8 on Automation-Bench, and 90.9 on GPQA Diamond. Independent scoring is thinner but exists — Artificial Analysis, read directly today, puts DeepSeek V4.1 Flash at 40 on its Intelligence Index, ranked 6th of 113 in its comparison class. One line from that same page matters more for a coding-agent reader than the rank does: Artificial Analysis labels the model very verbose, having consumed 250M output tokens to complete its Intelligence Index run against a 140M median. We will come back to that.
Two naming facts save real debugging time. DeepSeek's canonical API model id is now code>deepseek-flash/code>. The older strings code>deepseek-v4-flash/code> and code>deepseek-v4-flash-vision-exp/code> are still accepted, but the models behind them have been retired and requests are served by V4.1 Flash at the Flash price. Separately, DeepSeek V4 Pro is not gone: DeepSeek's documentation states the V4 Pro API service continues past 2026-09-14 with unchanged billing. If you were told the flagship was switched off, that is not what the vendor's own docs say.

OpenCode: two routes in, and the id to actually type
There are two ways to put DeepSeek V4.1 Flash behind OpenCode, and they have different economics.
The subscription route is OpenCode Go, a $10/month plan that OpenCode describes as a curated set of open coding models it has tested and benchmarked. Setup is four steps and there is no config file to hand-edit: sign in to OpenCode Zen, subscribe to Go, copy the API key, then run code>/connect/code> in the TUI, choose OpenCode Go, and paste the key. After that code>/models/code> lists what is available. Model references in configuration take the form code>opencode-go/<model-id>/code>.
Which model id? Both. The Go gateway's own model list, fetched today from code>https://opencode.ai/zen/go/v1/models/code>, returns code>deepseek-flash/code> and code>deepseek-v4.1-flash/code> as two distinct entries, alongside the legacy code>deepseek-v4-flash/code> and code>deepseek-v4-pro/code>. Either of the first two resolves to the model you want; code>deepseek-flash/code> is the canonical name and the safer choice for anything you intend to keep running.
The direct route skips the subscription entirely: run code>/connect/code>, search for DeepSeek, and paste a DeepSeek platform key. You are then billed by DeepSeek at list price rather than drawing on a Go allowance. DeepSeek publishes list price as off-peak $0.15 per 1M input tokens and $0.60 per 1M output tokens, with cached reads at $0.003 per 1M — and exactly double those during peak hours. Both the OpenCode Go documentation and OrcaRouter's model page, read today, agree on those figures.
What OpenCode Go adds is the allowance structure, and here the numbers are worth reading carefully because they are the reason the deadline matters. OpenCode's own docs list DeepSeek V4.1 Flash with a monthly limit of $15, currently multiplied 4× to $60 under a promotion OpenCode marks "Ends Sep 20". Estimated request counts are published in two columns: 6,500 per 5 hours / 16,250 per week / 32,500 per month at the standard rate, or 26,000 / 65,000 / 130,000 with the 4× multiplier applied. The shape of the allowance is consistent across models — the 5-hour limit is 20% of the monthly figure, the weekly limit is 50%, and the monthly limit is the whole thing.
Those are OpenCode's published numbers, read from its Go documentation today. The promotion is theirs to end, and it is dated.

Command Code: still live, and a bug report worth knowing
Command Code's own catalogue still lists the model as of today, under the id code>deepseek-v4-1-flash/code>, with a 1M-token context window and the same pass-through economics: $0.15 / $0.60 off-peak, $0.30 / $1.20 peak, $0.003 cached read. The matching prices across two independent harnesses are not a coincidence — both are passing through DeepSeek's list price rather than setting their own.
The mechanics are straightforward. Install with code>npm i -g command-code/code>, run it from your project directory, authenticate with code>/login/code>, and use code>/connect/code> if you want to bring your own provider key. code>/model <id>/code> applies a model change directly, while a bare code>/model/code> opens a picker, and code>/effort/code> sets reasoning effort for the current model — the flag that matters most here.
One community bug report is worth filing in your head before you debug the wrong layer. An open issue against a third-party routing layer describes a reasoning-replay cache that never engages for the code>command-code/deepseek-v4.1-flash/code> id, because the pattern the routing layer matches on expects a code>v4./code> or code>v4-/code> segment and the string is code>v4.1/code>. The reported symptom is upstream 400 errors complaining that reasoning content produced in thinking mode must be passed back to the API. That is a community bug report about a client-side matcher, not vendor guidance and not a problem with the model — but it is exactly the kind of thing that looks like a model fault at 2am.
Claude Code, Codex, and the clients OpenCode itself has validated
OpenCode Go is not OpenCode-only, and their documentation says so explicitly: it is designed for OpenCode and other coding agents that produce similar request patterns, with a published list of clients validated to work. That list currently names Hermes, Claude Code, Codex, ZCode and Pi — and for Claude Code the note is that it "recognizes its native session header. No custom-header wrapper is needed." Two requirements come with it: identify your client with its own user agent rather than a generic SDK name, and send a stable session identifier in the code>x-opencode-session/code> header on each conversation, which is what lets their routing and prompt caching work. For Hermes the validated build matters — the header fix merged after v0.21.0, so that release alone does not carry it.
There is also a route with no subscription attached at all, using DeepSeek's Anthropic-compatible endpoint directly. Set code>ANTHROPIC_BASE_URL/code> to code>https://api.deepseek.com/anthropic/code>, code>ANTHROPIC_AUTH_TOKEN/code> to your DeepSeek key, and point the model variables at the model. Claude model names are remapped on the way in: anything beginning code>claude-opus/code> goes to DeepSeek V4 Pro and is billed at the V4 Pro price, while code>claude-sonnet/code> and code>claude-haiku/code> names go to the Flash model. The code>[1m]/code> suffix on the model string requests the million-token context variant. DeepSeek's own guide for this setup also sets code>CLAUDE_CODE_EFFORT_LEVEL=max/code> and pins the auto-compact window to 786432 tokens.
That last variable is where a community fix lives. A workaround proxy exists because recent Claude Code builds send code>thinking: {"type": "disabled"}/code> on subagent requests while code>CLAUDE_CODE_EFFORT_LEVEL=max/code> adds a reasoning-effort parameter, and DeepSeek's Anthropic-format endpoint rejects that pairing with a message about thinking options not being disableable when reasoning effort is set. The reported workaround is narrow — strip the effort parameter from subagent requests only, leaving the main agent untouched. Treat it as a community finding about a client/server contract mismatch, and expect the exact version boundary to move.
The effort dial: 1–100, and why "low" means 50
Everything in this section is a community finding rather than vendor guidance. DeepSeek does not publish a preset-to-number mapping we could verify, and the numbers below come from practitioner write-ups and third-party harness documentation. They are consistent enough across sources to be useful, and unconfirmed enough that you should test them on your own work.
DeepSeek V4.1 Flash is trained with a continuous reasoning-effort scalar from 1 to 100. It is not a token cap — it moves where the model sits on a curve the training process learned, where lower effort applies more pressure to be concise and higher effort makes additional reasoning cheaper. Three public presets map onto that scale:
• Low — 50, the shortest path, and the preset that gives the control its reputation for cheapness.
• High — 75, and the level most community sources describe as the sane ceiling for agent work.
• Max — 100, where the penalty on reasoning length is removed entirely.
What the dial buys is real but sharply diminishing. One community benchmark sweep reported that raising effort from 25 to 100 moved Terminal-Bench 2.1 from 82.4 to 90.6 while roughly 2.5×-ing output tokens overall. Reports converge on effort somewhere between 60 and 80 capturing most of the available accuracy for under half the token budget, with max adding a further 1.6–1.8× to agent trajectories for a marginal gain.
The default is the part nobody agrees on. Some harness documentation and practitioner reports say an unset effort resolves to high; others describe the server default as simply unknown. What is documented rather than debated is that two harness integrations were found failing to send the parameter at all — the OpenCode Go provider profile and a native DeepSeek profile both skipped emitting code>reasoning_effort/code> for the code>deepseek-flash/code> slug, because their match guard expected a code>deepseek-v…/code> prefix the canonical id does not have. In both cases the user's chosen setting was silently replaced by the provider default. If your client shows an effort control, that is not evidence it is on the wire. Log one request body and look.
One more quirk from the same reports: the OpenCode Go endpoint accepts code>low/code>, code>medium/code>, code>high/code> and code>max/code>, but rejects an integer effort — a value of 80 has been reported returning HTTP 400. The 1–100 dial exists in the model, but it is not exposed as a raw number everywhere, so "set effort to 65" may not be expressible in your client.
The overthinking failure mode, and what actually fixes it
This is the failure mode that decides whether you keep the model in your loop. Community reports describe DeepSeek V4.1 Flash continuing to reason after the work is done: re-arguing a point it has already answered correctly, narrating the correction of its own wrong assumptions, and producing long reasoning chains with low information density. One practitioner reported reasoning output continuing for over an hour inside a coding CLI session. Another reported being unable to get it to finish a long benchmark run at all.
A sourcing note on that second report, because it is the kind of claim that gets laundered. It comes from a community thread on running the model reliably, and Reddit blocks our fetcher, so we could not read the thread directly — we are relaying the report rather than quoting a page we opened. What we could verify independently is the shape of the problem, and there the outside evidence is unusually clean: Artificial Analysis, on its own page, flags the model as very verbose, consuming 250M output tokens to complete a run its median comparison model finishes in 140M. That is roughly 1.8×, measured by a third party, on a fixed task set. The forum reports and the independent metric are describing the same behaviour.
The mitigations that come out of those reports, all community-sourced:
• Pin effort at high or below and refuse to let it ratchet. The most explicit fix seen in the wild is a routing plugin written specifically to stop effort escalation: turn depth contributes nothing to the escalation score, only a failed tool result or an identical retry counts, escalation is capped, and max is opt-in and demoted by default. If your harness lets an agent raise its own effort as a run gets longer, that is the mechanism to switch off.
• Do not run max as a default. Multiple practitioners report max spinning rather than converging on routine work and switching back to high resolving it.
• Cap code>max_tokens/code> on interactive paths. A 384,000-token output ceiling is a limit, not a target, and a loop that will not terminate is expensive at that ceiling.
• Verify the parameter is being sent. Given that two harnesses were found dropping it silently, "I set it to high" and "high reached the API" are different claims.
• The harness matters more than you would expect. Practitioners running the same weights report sharply different behaviour across shells — the same model that argues with itself and buries the signal in one harness drawing none of those complaints in another. That is a community observation about harness behaviour, not a vendor claim about the model, but it is the most repeated piece of advice in the reports.
What a coding loop actually costs at these rates
DeepSeek's list price is $0.15 per 1M input tokens and $0.60 per 1M output tokens off-peak — output is four times input, which is the first thing to internalise about an agent that generates reasoning as well as code.
Take a realistic agent turn: 60,000 tokens of context (system prompt, tool schemas, a repo slice, conversation history) in, and 3,000 tokens of reasoning plus a patch out. That is 60,000 × $0.15/1M = $0.009 in, plus 3,000 × $0.60/1M = $0.0018 out, so about 1.1 cents a turn. Two hundred such turns in a working day is roughly $2.16, or about $47 across a month of weekdays at the off-peak rate. That is the arithmetic that makes a $15 monthly allowance with a 4× promo on top look generous — and the arithmetic that makes verbosity the thing to watch.
Because here is the leverage hiding in that Artificial Analysis number. The model using 1.8× the median output tokens on a fixed task set means an output-bound loop costs 1.8× what the token price alone suggests. And the effort dial is the control for exactly that. Community reports put the move from max down to high at roughly halving output tokens — a far larger swing than anything the peak/off-peak schedule will do to you. The effort setting is the big lever; the schedule is the free one.
Which is worth knowing before you schedule anything: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Monday through Friday, and everything else including weekends is off-peak. A European team working 09:00–18:00 CET never touches peak. A team in Beijing working the same local hours hits peak from 09:00–12:00 and 14:00–18:00 — seven of nine working hours at double price. Same model, same code, double the bill, decided entirely by timezone.
The other free saving is the cached read rate, $0.003 per 1M against $0.15 for fresh input — one fiftieth. A coding agent's prompt is mostly stable prefix: system instructions, tool definitions, the parts of the repo that are not changing. Keep the stable material at the front and let the variable content trail it, and the provider-side cache does the rest. Whether discounted cached input applies, and on what conditions, is set by DeepSeek rather than by the client, so confirm it in current documentation before you build a budget on it — our own model page for code>deepseek/deepseek-v4.1-flash/code> says the same thing, that the cached-input rate follows the provider's terms.
If you would rather not add a second subscription to evaluate the model, the same id is available through one OpenAI-compatible endpoint in front of our whole catalogue, at the provider's rate with 0% markup — so a price change on DeepSeek's side is live here the same day rather than at the next repricing. That matters most for exactly the situation this article describes: a model with a documented tendency to keep going, on a path you have not finished evaluating. A fallback chain means a turn that goes wrong lands on another model before the response starts, instead of failing the request.
552B, 748B, or 763B — the size question is genuinely open
Do not repeat a parameter count for this model as settled, because three different numbers are in circulation and none of them is simply wrong.
DeepSeek's own model card describes a "552B backbone parameters" model, and that is the vendor-reported figure — it is also the number carried in the spec panel on our own model page. The same card separately lists an "Engram conditional memory" module of 196B parameters, "sparsely accessed via token-based lookup." Add the two and you get 748B, which is the arithmetic the community analysis landed on within hours of release. And the file metadata on the same model repository lists a model size of 763B parameters, which is a third figure again.
The apparent dispute is definitional rather than a contradiction. Looked-up Engram parameters cost memory but almost no arithmetic per token, while computed backbone parameters cost time on every token — which is exactly why the vendor reports them separately and why comparing the headline numbers of two models with different architectures tells you very little.
What is not seriously contested is the active parameter count: roughly 8B per token during prefill and 16B during decode, which is the number that actually governs inference cost. The weights are open under an MIT licence if you want to check any of this yourself. Treat 552B as vendor-reported, 748B as a credible community total, and any claim that the size question is closed as premature.

What to do before the 20th
If you are going to try DeepSeek V4.1 Flash in a coding agent, the ordering that wastes least time is: pick the harness first, then pin effort, then measure.
On the harness, the honest summary of today's verification is that all three routes work and they differ in what they ask of you. OpenCode Go is a $10/month subscription with a promotional multiplier that expires on 20 September, and the setup is code>/connect/code> plus code>/models/code> with no file to edit. Command Code lists the model live in its own catalogue, switches with code>/model <id>/code>, and exposes effort as a first-class command. Claude Code reaches it either through OpenCode Go's validated-clients path — where it needs no custom header wrapper — or directly against DeepSeek's Anthropic-format endpoint, where the subagent effort conflict is the known rough edge.
On effort, set it explicitly and set it low-ish: high, or the model default if high is what that turns out to be, and not max. Then confirm it left your machine, because two harnesses were found dropping it.
On measurement, watch output tokens rather than wall-clock. The verbosity is the cost, the effort dial is the control, and the peak schedule is a timezone accident you can avoid for free.
And on the size claims you will see quoted at you this week — 552B, 748B, 763B — the useful response is that the vendor reports its backbone, the community adds the memory module, and the repository metadata says something else again. Anyone presenting one of those as the settled answer has picked a number rather than checked one.
