
DeepSeek Harness, Explained: The Agent Runtime Behind "Model + Harness = Agent"
- OrcaNEWOrca: OrcaCyber Zero 1.52026-10-10$3.00 / $7.50 per 1M tokens · 83 tok/s
- openaiNEWOpenAI: GPT-6.1 Sol2026-09-2952Intelligence
- anthropicNEWAnthropic: Claude Sonnet 5.52026-09-2856Intelligence
- typesafeTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 122 tok/s
- OpenAIOpenAI: GPT-6 Luna2026-09-2238Intelligence
- OpenAIOpenAI: GPT-6 Sol2026-09-2248Intelligence
- AnthropicAnthropic: Claude Opus 5.52026-09-2258Intelligence
- xAIGrok 4.72026-09-2146Intelligence
- OrcaOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $7.50 per 1M tokens · 53 tok/s
- OrcaOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 423 tok/s
- DeepSeekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- OpenAIOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- GoogleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- AlibabaQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- AnthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- TencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens · 60 tok/s
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 354 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 231 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
DeepSeek Harness — dsh for short — is not a model. It is the open-source agent runtime DeepSeek released under the MIT license on August 13, 2026, and it is still the fastest thing DeepSeek has ever shipped by adoption: press coverage reported roughly 50,000 GitHub stars in its first 12 hours, and the repository stood at 242,211 stars with 29,059 forks when we read the GitHub API today. What it runs is DeepSeek V4.1 Flash, the model DeepSeek's API serves as deepseek-flash, or the larger DeepSeek V4 Pro when you point it there. Its job is the engineering layer between a language model and a finished task: file editing, terminal execution, web search, context management, task planning, and calling other agents. The formula DeepSeek's own team uses is Model + Harness = Agent — the model reasons, the harness does everything else.
That distinction is the point of this article. People searching "deepseek harness" are usually expecting a model review and landing on launch coverage instead. So here is the answer up front: DSH is not a model you call through an API. It is a program you install and run — the npm package is @deepseek-ai/dsh, whose latest tag now resolves to 0.2.0-rc.2 — three release lines past the 0.1.x build this page was first written against, and past the version chip on the card above — and it connects whatever model you give it (DeepSeek V4.1 Flash by default) to your filesystem, your terminal, the web, and other agents. It is MIT-licensed, built on a plugin framework called Cordis, and it is explicitly a developer preview that warns of breaking changes. The packaged desktop app DeepSeek announced on October 2 — macOS and Windows builds, with Linux users pointed at the npm package — is the newest surface on top of it. Below is what it actually does, how to run it today, and where it falls short.
What a harness actually does (and why the term is confusing)
A harness is the scaffolding that lets a raw model do work in the real world. A model alone can only produce text. Give it a harness and it can read and write files, run commands in your shell, search the web, keep a plan, call tools, react to errors, and judge when a task is finished. DeepSeek's public formula — repeated in launch coverage and its own job postings — is "Model + Harness = Agent": the model supplies reasoning, the harness supplies everything else.
This is why the release matters beyond one product. DeepSeek reported its own agentic benchmark scores for its last two models from inside this harness: the DeepSeek V4 Flash API documentation states its Terminal-Bench 2.1 result of 82.7 was produced on "DeepSeek Harness minimal mode." The harness is the reference environment those numbers came from, and it is now public for anyone to reproduce — or contradict.
Everything is a plugin: the Cordis architecture
DSH's core design, stated in its own README, is that everything is a plugin. The agent loop, model adapters, tools, skills, session storage, sandbox, scheduling, and even the web UI are all Cordis plugins. There is no privileged core to patch: you extend the harness by mounting a plugin next to the others, and every registration is an effect that unwinds when the plugin unloads. The plugin engine is Cordis, a meta-framework whose design is described in the Peking University–DeepSeek paper "A Programming Paradigm for Spatiotemporal Composability."
The practical consequence is a sharp separation of concerns: plugins add capabilities; presets decide which capabilities a given agent can see. The monorepo now carries 55 packages under packages/ — core, llm, mcp, sandbox, context, plan, goal, workspace, and more — and launch coverage counted more than 100 first-party plugins, with a plugin store already reserved and 288 community plugin repositories collected on GitHub within 24 hours of the release. That was six weeks ago. The dsh-plugin topic DeepSeek uses for plugin discoverability now returns more than 17,000 repositories in GitHub's search, which we read on October 2, 2026 — the ecosystem kept growing even as the harness itself stayed on release candidates.
The four presets — and when to use which
DSH ships with four agent presets, each loading a different plugin set. From the Web UI's own preset menu:
• Standard — a full coding agent: file editing, shell, file and web search, skills, planning, goals, subagents, and workflows. The right default for most work.
• Code — called PTC, "programmatic tool calling," in Chinese press coverage — everything in Standard, plus a Code Mode SDK that lets the model write TypeScript programs to compose multi-step tool calls in one program. More powerful, and it can multiply side effects and tool calls.
• Minimal — just persistent bash plus str_replace_editor. Built for reproducing benchmark runs; this is the preset DeepSeek used for the V4 Flash Terminal-Bench 82.7 figure. Minimal is not harmless — shell access is still shell access.
• Creator — the fourth preset directory is literally named cordis — Standard plus runtime inspection, in-memory Cordis plugin experiments, and preset authoring. The agent can install, replace, or destroy plugins mid-conversation. An engineering environment, not a production approval layer.

Which one should you pick? If you are reproducing a benchmark, Minimal. If you want the agent to write its own multi-step orchestration, Code. If you are building or testing a plugin, Creator. For everything else, start with Standard — broad authority means you also want a tight workspace and permission policy.
Trajectory: the feature reviewers call "DevTools for agents"
The most-praised feature in hands-on reviews is Trajectory. Every run writes to an append-only session log — zstd-compressed JSONL with a uniform event envelope of type, sequence, time, and data — recording everything the model saw: system prompts, reasoning, tool calls and results, subagent scheduling, context injections. The Trajectory view supports inspection by source, resume, fork, search, and replay. Forking is the standout move: you can change one instruction and replay the run side by side with the original instead of throwing the trace away.
Concretely, in a hands-on test published August 15, a single file-writing task produced 61 session events, and a trivial "reply PONG" task still consumed about 13,467 input tokens on its first request — overhead from default system prompts, tool schemas, auto-injected repo rules, and skill summaries. That is the honest cost of a harness that logs everything: you pay tokens for the scaffolding, and you can see exactly what you paid for.

MCP, plugins, and talking to other agents
MCP is where the hype outruns the current build. The monorepo includes an mcp package and the CLI ships a dsh-mcp-client dependency, but no MCP server is enabled by default, and the architecture routes tool integration through Cordis plugins rather than treating MCP as a first-class citizen. You can connect MCP servers as a tool source; it is just not the default path yet.
What is more surprising is what the subagent providers list shows. DSH can spawn other agents as subagents — including, per the capability-seams documentation, a Codex provider and a Claude Code provider. That is an unusual position for a competitor to take: DSH is a harness that can delegate work to the very products it is measured against.
The plugin story is the real platform bet. DeepSeek has reserved a plugin store, community repositories are tagged for discoverability, and third-party in-app stores already exist. The 0.2.0-rc.2 release notes spend most of their length on plugin installation guidance — distinguishing installed, incompatible and bundled plugins — which is where the engineering effort is still going. The long game is that the harness, not the model, becomes the ecosystem, which is why the pricing discussion around DSH matters more than any single feature.
How to run it today
You need Node.js 22.19 or newer — the package manifest requires ^22.19.0 or >=24 — and then one command:
npx @deepseek-ai/dsh web
That starts the Web UI, served at http://127.0.0.1:3080 by default. In the UI: open Settings → Models, paste a DeepSeek API key, then choose a workspace — the session composer stays locked until you do. The default model is DeepSeek V4.1 Flash, and the default permission policy is workspace-write with approval prompts for everything else.
There are now five entry points: the Web UI, the packaged desktop app, headless one-shot runs, a Python SDK that drives DSH over JSON-RPC stdio, and --dump-config, which prints the assembled plugin tree (the web profile assembled 129 lines of plugin config; headless, 81). The desktop app is the one that changed this week. DeepSeek published builds for macOS and Windows — Intel Macs included, where our September note recorded the Intel build 404ing — and every desktop feed reports version 0.2.0-rc.2 with a release date of September 29. The 0.2 notes also bundle the dsh command into the desktop app for plugin management, so that route no longer needs a separate Node or pnpm install. DeepSeek's own announcement describes plugins and Workspaces as supported out of the box in 0.2 — a vendor claim about a release candidate, though the plugin pages and the workspace store the app shares with the CLI are visible in the public repository. Linux is the exception DeepSeek's own announcement names: the packaging script's supported target set is mac-arm64, mac-x64 and win-x64, no Linux artifact exists under DeepSeek's download host, and Linux users are told to run npx @deepseek-ai/dsh web. From source the path is unchanged: git clone, pnpm install, pnpm run build, then pnpm dsh web.
The honest part: preview-grade, and where DSH is the wrong tool
DSH is version 0.2.0-rc.2 — not 1.0, and not even a stable 0.x: all 29 versions on npm and all 24 GitHub releases are prereleases, and the README is explicit: "THERE WILL BE COMPATIBILITY-BREAKING CHANGES." Reviews from the launch period catalogue the rough edges — silent failures under MSYS2 on Windows, hot-reload cache staleness, a roughly 200ms batched persistence delay that lags the UI progress display, and community criticism that GitHub Issues is disabled on a repository positioned as public infrastructure. The default product experience also still trails the polished coding agents it is measured against; one hands-on review concluded the open-source release is an agent runtime scaffold, not yet a finished product.
Where is DSH the wrong tool? Four situations. First, if you want turnkey polish — today's Claude Code and Codex are more finished, and DSH is for people who want the scaffold and are willing to finish it. Second, if you cannot audit what you run: plugin provenance, config diffs, and credential boundaries are all self-supplied, and production use requires governance you provide. Third, if you are benchmarking models: you cannot run one model in Standard mode and another in Minimal and call the difference a model result — you changed the harness, not only the model. Fourth, if you are cost-sensitive about token overhead: the always-on context can burn thousands of input tokens before the model does any real work.
What it means for DeepSeek V4.1 Flash and DeepSeek V4 Pro
DSH defaults to DeepSeek V4.1 Flash — the model DeepSeek's API serves as deepseek-flash — and the flagship DeepSeek V4 Pro build (DeepSeek-V4-Pro-0813) is the same model family the harness was built to run. Both are available today through ordinary API calls, no harness required, and both are hosted on OrcaRouter at DeepSeek's own list price, passed through at 0% markup: $0.15 input and $0.60 output per million tokens for DeepSeek V4.1 Flash, and $0.66/$1.98 for DeepSeek V4 Pro. Those are off-peak rates, which DeepSeek doubles between 01:00 and 04:00 and again between 06:00 and 10:00 UTC on weekdays. The screenshot below is an August snapshot of the DeepSeek V4 Flash 0731 model page, and the $0.15/$0.29 it prints is that release's price: DeepSeek's pricing page now exposes a single Flash entry, deepseek-flash = DeepSeek-V4.1-Flash, at $0.15/$0.60. Read the card as a dated capture, not as today's rate.

The honest boundary: OrcaRouter is a model router, not an agent runtime, so we do not host DSH — it is not a model. What a router genuinely answers here is the failover problem DSH raises. The harness is a fast-moving preview — 0.2.0-rc.2 today, with the desktop app on top of it younger still — and pointing production traffic at it means betting on vendor-reported agentic scores. Running the same DeepSeek models through a routing layer on one key lets you test DSH's native numbers while keeping automatic failover across vendors for the traffic that has to keep working. The harness is the thing to experiment with; the models it runs are the thing to route carefully.
Bottom line
DeepSeek Harness is the most consequential thing DeepSeek has open-sourced since the models themselves, because it changes what "DeepSeek" means: not just models, but the agent layer that runs them. For a reader deciding what to do: if you build agents, install the desktop app or the 0.2.0-rc.2 command line this week and run the reproduction test — the Terminal-Bench 82.7 DeepSeek published for its Flash model was produced inside this harness, and now you can check it yourself. But treat it as what it is: a developer preview with breaking changes ahead, a plugin system that rewards people who like owning their stack, and a harness that is still rougher than the polished alternatives. If you are on Linux, the packaged app is not the route for you yet — npm is. The models underneath are the reliable part, and those you can route with confidence today.
Compared in this article1
Detected from this article · Benchmarks: Artificial Analysis · updated daily
