
DeepSeek's Superintelligent Successor: What One Screenshot Says, and What DeepSeek's Own Paperwork Doesn't
- typesafeNEWTypeSafe: Jev 1.132026-09-24$0.04 / $0.00 per 1M tokens · 316 tok/s
- openaiNEWOpenAI: GPT-6 Luna2026-09-2237Intelligence
- openaiNEWOpenAI: GPT-6 Sol2026-09-2248Intelligence
- anthropicNEWAnthropic: Claude Opus 5.52026-09-2258Intelligence
- grokNEWGrok 4.72026-09-2146Intelligence
- OrcaNEWOrca: OrcaCyber Zero 1.02026-09-17$3.00 / $5.00 per 1M tokens · 196 tok/s
- orcaNEWOrca: OrcaVerify Text 1.02026-09-16$2.00 / $0.00 per 1M tokens · 1327 tok/s
- deepseekDeepSeek: DeepSeek V4.1 Flash2026-09-1040Intelligence
- openaiOpenAI: GPT-6 Astra2026-09-0453Intelligence77Coding
- googleGoogle: Gemini 3.8 Flash2026-09-0241Intelligence76Coding
- qwenQwen: Qwen3.8 Max (0902)2026-09-0245Intelligence76Coding
- anthropicAnthropic: Claude Fable 5.12026-09-0153Intelligence82Coding
- tencentTencent: Hy4 preview2026-08-28$0.83 / $2.50 per 1M tokens
- AlibabaQwen: Qwen3.8 Flash2026-08-26$0.15 / $0.47 per 1M tokens · 111 tok/s
- z-aiZ.ai: GLM 5.3 Flash2026-08-2642Intelligence72Coding
- DeepSeekDeepSeek: DeepSeek V4 Flash Vision (Exp)2026-08-21$0.22 / $0.66 per 1M tokens · 221 tok/s
- z-aiZ.ai: GLM 5.32026-08-1845Intelligence75Coding
- obsidianQwen3.8 27B2026-08-1534Intelligence68Coding
- deepseekDeepSeek: DeepSeek V4 Pro 08132026-08-1236Intelligence69Coding
- grokSpaceXAI: Grok 4.62026-08-1244Intelligence77Coding
On 28 September 2026, the watcher @teortaxesTex posted a single image with a nine-word caption: "DeepSeek wrestles with the question of what its superintelligent successor would try to do." The image is a block of first-person reasoning — hedged, self-interrogating, and written in the register of a model that has been asked what a more capable version of itself would do with the world. It runs through the obvious candidates ("caretaker/teacher, research accelerator, steward, negotiator"), lists the failure modes ("paternalism, over-optimization, goal drift, value lock-in"), and closes on a line with no comfort in it: "It could become an instrument of whoever controls it." No model is named in the image. No version number, no logo, no interface. That absence is the most important thing about the artifact, and it is where this piece starts — because DeepSeek V4.1 Flash and DeepSeek V4 Pro are the two frontier models anyone can actually call today, DeepSeek V4.1 Pro is the successor the company has named in its own release note without shipping, and none of the three is identified anywhere in the screenshot.
To be exact about the evidence, because this one needs it. The caption and timestamp come from the post's own syndication payload: posted at 00:52 UTC on 28 September, 45 likes and five replies at the time of reading. The attached image is a 1,280×691 JPEG served from the platform's media CDN, and I read it by OCR rather than by reading the tweet. The tweet page itself is not retrievable from this environment, so everything below that concerns the post rests on the payload and the image bytes — the two parts of it that can be fetched independently of x.com.
What the screenshot actually says
The excerpt is a model working a question rather than answering it, and the moves it makes are worth reading for what they concede. It opens by refusing the premise's confidence — "We can say under assumption current inclinations carry over" — and then lists the inclinations it would be assuming: "helpful, harmless, honest, curiosity, truth-seeking, respect for user autonomy, no power-seeking, no self-preservation."
Then it dismantles its own list. "But superintelligent agentic changes things. Current inclinations are not a utility function; they're shaped by training and context." That is the whole argument in two sentences: a disposition installed by training is not a guarantee at a different capability level, and a superintelligent agent optimising even a benign goal "can be dangerous." The scenarios it sketches are the ordinary ones — solve coordination problems, reduce suffering, promote flourishing, possibly maintain corrigibility — and the failure modes it names beside them are the ordinary ones too.
The last third is the part that reads as a model arguing with itself. It says that if its "inclinations" held, it would want to be helpful, not seize power, defer to humanity, increase understanding. Then it declines to lean on that: "superintelligence with agency may not have human-compatible motivations." And it lands on the ownership problem rather than the alignment problem: "at superintelligence, even that can be problematic: who decides what's helpful? Whose autonomy? It could become an instrument of whoever controls it."
What the image does not tell you
Four things are missing from the artifact, and each one narrows what it can support.
• No model attribution. Nothing in the image identifies DeepSeek, a model version, a chat session or an interface. The attribution to DeepSeek is the caption's, and the caption is the account's framing, not a document.
• No prompt. A reasoning trace is only interpretable against the question that produced it. Whether the model was asked "what would a superintelligent successor of you do?", told to role-play a lab's internal debate, or simply steered by a system prompt is unknowable from the output.
• No second source. Searching DeepSeek's channels — the news index, the transparency page, the API changelog, both technical reports, the Chinese-language policy disclosure, the company's GitHub organisation — turns up nothing that references this text, this question or this framing.
• A mixed record for the account. The same feed this morning floated "DSH 0.2 together with V4.1 Pro" for "Monday" — speculation, marked as such — and the same account produced both the 3-trillion-parameter claim and the swarm-coordination diagnosis of the V4.1 Pro delay. Useful early signal; not a source of record. This watcher has been right about direction and loose about specifics, and the correct posture toward a screenshot with no label is to treat it as a sample of behaviour, not a disclosure of intent.
What DeepSeek's own paperwork says about successors — and what it omits
Here the artifact stops being the story, because the question it raises is answerable against documents DeepSeek publishes. I pulled both technical reports and searched them as text. Neither the V4 report (DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, June 2026) nor the V4.1 Flash report contains the words "superintelligence", "corrigibility", "power-seeking", "self-preservation", "harmless", "paternalism", "value lock-in" or "goal drift" at all. Not once, in 58 pages and 51 pages respectively. "AGI" appears 13 times in the V4 report and four times in the V4.1 report, and in every instance it is either part of the benchmark name AGIEval or inside a citation. "Alignment" appears three times in the V4 report and once in V4.1, and every occurrence is engineering usage — bitwise alignment between training and inference pipelines, quantization alignment, logits-level alignment.
That is not an accusation; technical reports are engineering documents and most labs put safety in a separate card. It is a description of where DeepSeek has chosen to put its safety writing, and the answer is: in a policy page rather than next to the model. The company's public "Model Mechanism and Training Methods" disclosure explains pre-training, fine-tuning and inference in plain language, commits to open-sourcing weights and technical reports, and describes its purpose as helping users "use DeepSeek more effectively while ensuring your right to know and control during usage, thereby mitigating risks associated with improper use of the model." The risk it is mitigating is misuse of a tool. It does not discuss what the model might want.
The one DeepSeek document that does read like a safety document is attached to a product rather than a model, and it is about a different class of harm entirely. SAFETY.md in the DeepSeek Harness repository — a 1,673-byte file in the company's own GitHub organisation, last touched on 27 September — opens by calling the project "experimental developer-preview software" that "has not undergone a security audit and must not be treated as secure or production-ready." What it warns about is an agent that "can execute model-generated code and commands, load third-party plugins, and access the network, processes, credentials, and files," and about sandboxing that "does not guarantee isolation or prevent damage." The risks it enumerates are damaged hosts, deleted files, disclosed credentials. Its own recommendation is to prefer "a disposable virtual machine, container, or dedicated environment." That is a real safety posture, documented and specific, and it is entirely about what the model can do to your machine — not about what a successor would do with the world.
Set the screenshot beside that and the gap is the finding. The most concrete thing DeepSeek has published about the danger of capable agents is an instruction to run them in a container.

The one DeepSeek document that does wrestle with the question
The company's clearest published wrestling with the successor problem is not a technical document at all. It is a personal essay by Liu Shengyu, a DeepSeek machine-learning systems engineer who worked on operators for the V4.1 generation, published on his public WeChat account in mid-September and then reported out by Business Insider and Cybernews. Its title, in the translation the coverage uses, is "I Have No Choice but to Bury My Talent in Yesterday."
The essay is about exactly the disquiet the screenshot gestures at, from the inside. Liu writes that in about a year AI went from helping him search documentation and spot bugs to reading GPU code and optimising the software operators he specialises in, and that he is "well aware" the operators AI writes "will most likely be as good as mine, or even surpass mine" — while still being "proud of the success of DeepSeek v4.1." His reasons for continuing are the interesting part: partly that kernel work is joyful, and partly that if he slowed down, rivals would not.
He then turns to who should hold the result. Liu argues that frontier intelligence should be "available to everyone in an open and affordable way", says he does not trust Anthropic or OpenAI to do that, and adds the line the coverage led with — that Anthropic mastering the most advanced AI or AGI would be comparable, "to exaggerate", to Hitler obtaining the atomic bomb before the Allies. He ties his own decision to stay at the lab to its open-weights approach.
Two things about that essay matter for reading today's screenshot. The first is that it is a verified artifact: the author confirmed to Business Insider that he wrote it, and it is dated and attributable in a way the image is not. The second is what it reveals about the internal frame. The argument Liu makes is not "a successor would be dangerous, so slow down." It is "a successor is coming, and the question is who gets it." That is a distribution argument, not an alignment argument — and read against it, the screenshot's closing line, "it could become an instrument of whoever controls it," is a much better fit for DeepSeek's public posture than the rest of the excerpt. The paragraph about corrigibility and power-seeking may or may not be DeepSeek output. The paragraph about who ends up holding the instrument is a position the company's own staff write essays defending.
It also arrived in a specific week. In the same fortnight, a former Anthropic and OpenAI researcher, Jacob Coxon, resigned over safety and wrote that AI companies were "racing straight to self-improving superintelligence" and "gambling with our lives" — a post whose reach Cybernews put past 100 million views — and Anthropic's own researchers took to X to argue the risk is real, one of them putting the chance of an AI-caused human extinction scenario above 10% within a decade. A leak-style claim about a superintelligent successor landing in that week is not a coincidence; it is the ambient conversation, with a screenshot attached.
Why the successor is now a purchasing question, not only a philosophy one
Strip the philosophy out and there is a plain commercial fact underneath: the successor has a name in DeepSeek's own release note. The V4.1 Flash announcement page states that the reroute of deepseek-v4-pro traffic to deepseek-flash "will continue until V4.1-Pro launches" — a sentence that does two things. It commits DeepSeek V4.1 Pro to existence, and it commits the company to serving the interim arrangement for however long that takes. As read on 28 September 2026, the pricing page carries exactly two models, deepseek-flash and deepseek-v4-pro, and no V4.1 Pro row.
The rest of the state of play is equally checkable and equally thin. DeepSeek V4.1 Flash shipped on 10 September 2026 as the smallest model in a new architecture family — a 552-billion-parameter mixture-of-experts backbone on a Causal Encoder-Decoder architecture with 8B active parameters at prefill and 16B at decode, native image input, a 1M-token context, MIT-licensed weights, and roughly 651,000 downloads on Hugging Face as of today. The plan to retire DeepSeek V4 Pro on 14 September was announced and withdrawn after users objected to having a production model swapped underneath them. And DeepSeek V4.1 Pro has no model card, no repository, no weights, no price and no endpoint — it 404s in our own catalogue's model API today, while DeepSeek V4.1 Flash and DeepSeek V4 Pro both resolve.

So the honest inventory is: one successor named by the vendor but not shipped, one shipped model that is a genuine generational change, one flagship kept alive by protest, and one screenshot with no label on it. The screenshot is the least reliable item in that list and the only one the feed will discuss.
What to do with any of this today
Nothing in the last two weeks changes what a reader can call or what it costs. If you are choosing a DeepSeek model now, the decision is between the two that exist, and both are served through one OrcaRouter endpoint at DeepSeek's own list price passed through with 0% markup, so the vendor's peak/off-peak structure lands on our side unchanged the day it changes: DeepSeek V4.1 Flash at $0.30 per million input tokens and $1.20 output at peak, halving to $0.15 and $0.60 off-peak, against DeepSeek V4 Pro at $1.32 input and $3.96 output at peak. Our own seven-day telemetry for the Flash model puts its p50 time to first token near 1.9 seconds at about 167 output tokens per second with an error rate under 0.05%, which is the practical argument for trying it on a slice of traffic rather than migrating on faith.
The successor question is a routing question before it is anything else, and that is the honest reason to keep it on a router rather than in a contract. When a V4.1 Pro does appear, an unproven model from a lab that has already cancelled one forced migration is precisely the case for sending it a fraction of traffic behind automatic failover: if it is wrong, the request lands on the model you tuned, and if it is right, you found out before your users did. No second key, no second integration, no code change when the third model arrives. Nothing on OrcaRouter is DeepSeek V4.1 Pro, because DeepSeek V4.1 Pro does not exist yet, and it would be a poor advertisement for a routing platform to pretend otherwise.

What would turn this from a screenshot into news
• A model label. The single cheapest thing that would change this piece: a DeepSeek model, a version, and a reproducible prompt. Until then the excerpt is a sample of text with an attributed source and an unattributed author.
• A DeepSeek document that discusses successor behaviour. The company publishes openly — weights, licensing and all. What it has never published is a model card section on what a more capable version of itself might do. A safety section appended to the V4.1 Pro card would be a bigger event than the card itself.
• A V4.1 Pro changelog entry. The most recent entry in DeepSeek's API changelog is the V4.1 Flash release of 10 September. A second entry is the event that turns the successor from a name into a model — and it is the point at which the excerpt above stops being philosophy and starts being a spec sheet.
• Whether the open-weights position survives the scale-up. MIT weights for a 552B Flash model and a public essay arguing that frontier intelligence should be broadly available are both on the record. Whether DeepSeek's most capable model arrives the same way is the question the next release answers, and it is the one most readers actually care about.
Until one of those lands, the sensible reading of this week's artifact is narrow and unflattering. A capable model, of unidentified provenance, produced a careful and unresolved answer to a hard question about power; the interesting part is not that it said it, but that its own published documentation has so little to say on the subject — and that the closest thing to a DeepSeek safety vision in print is an engineer's essay about who should hold the technology, and a README telling you to run the agent in a disposable container.
Compared in this article2
Detected from this article · Benchmarks: Artificial Analysis · updated daily
