DeepSeek V4 Flash

deepseek/deepseek-v4-flash
ToolsJSONReasoning
by DeepSeek · 2026-04-24

DeepSeek V4 Flash efficient MoE — 284B total / 13B active params, 1M context, optimized for fast everyday workloads.

ctx1M tokens
Max output384K
Inputtext
p50 TTFT1.94 s
INPUT$0.22/ 1M tokens
OUTPUT$0.66/ 1M tokens
p50 TTFT1.94 s7d
p95 TTFT10.00 s7d
TRAFFIC11870.3Mtokens / 7d

DeepSeek V4 Flash is a large language model from the Chinese AI company DeepSeek. It processes text inputs only and is designed for scenarios that demand a large context window (1,048,576 tokens) and…

What is DeepSeek V4 Flash?

Who should use DeepSeek V4 Flash?

What input modalities does DeepSeek V4 Flash support?

Code samples

Call from any SDK

OpenAI-compatible — keep the SDK you already use

  • OpenAI SDKhttps://api.orcarouter.ai/v1
import os

from openai import OpenAI

client = OpenAI(
    base_url="https://api.orcarouter.ai/v1",
    api_key=os.environ["ORCAROUTER_API_KEY"],
)

response = client.chat.completions.create(
    model="deepseek/deepseek-v4-flash",
    messages=[{"role": "user", "content": "Hello"}],
)
print(response.choices[0].message.content)

Supported parameters

  • include_reasoning
  • logprobs
  • max_tokens
  • reasoning
  • response_format
  • stop
  • stream
  • stream_options
  • temperature
  • thinking
  • tool_choice
  • tools
  • top_logprobs
  • top_p
  • user_id

Pricing

Pricing
Input / 1M tokens · Off-peak$0.220
Output / 1M tokens · Off-peak$0.660
Cache read / 1M · Off-peak$0.0070
Peak hours01:00–04:00, 06:00–10:00 ×2 (UTC)
Input / 1M tokens · ×2$0.440
Output / 1M tokens · ×2$1.32
Cache read / 1M · ×2$0.014
CurrencyUSD

Cost calculator

Tokens / month10MM
Input share70%%
Estimated / month $3.52 · With prompt caching $2.77

Estimate based on list price

Token & cost estimator

Input tokens: 20Cost per request: $0.000334

Estimate only — actual token counts depend on the provider's tokenizer.

Performance

p50 TTFT
1.94 s
Output speed
170 tok/s
p95 TTFT
10.00 s
Error rate
17.6%

Public benchmarks

69.1
AA Coding
Better than 79% of models compared
#28 of 138
34.3
AA Intelligence
Better than 69% of models compared
#44 of 147
50.0
AA Math
Better than 26% of models compared
#61 of 82
GPQA Diamond
90.8
Humanity's Last Exam
38.6
IFBench
79.2
Long-Context Recall
79.7
MMLU-Pro
57.0 index
SciCode
50.3
tau_banking
39.4
TerminalBench Hard
35.6
terminalbench_v2_1
78.7
terminalbench_v4_0
12.1
τ²-Bench
95.0
Source: artificialanalysis.ai

Community buzz

What developers are saying this week

Hacker News1 mentions · 7ddown 1 vs the previous week

How it compares

DeepSeek V4 FlashDeepSeek V4 ProDeepSeek V4.1 FlashDeepSeek V4 Flash Vision (Exp)
Input $/M$0.22$0.66$0.15$0.22
Output $/M$0.66$1.98$0.60$0.66
Context1.0M1.0M1.0M1.0M
Quality7/108/108/107/10
Compare side-by-sideCompare side-by-sideCompare side-by-sideCompare side-by-side

FAQ

How much does DeepSeek V4 Flash cost on OrcaRouter?
Input tokens cost $0.14 per 1 million tokens, and output tokens cost $0.28 per 1 million tokens. OrcaRouter charges exactly the provider rate with zero markup.
What is the context window of DeepSeek V4 Flash?
The context window is 1,048,576 tokens (1 million tokens). The maximum output per request is 384,000 tokens.
What are the main strengths of DeepSeek V4 Flash?
Its strengths are the very large context window and high output token limit, combined with a low price point and a strong τ²-Bench score of 95.0, indicating good reasoning and tool-use capabilities.
How does DeepSeek V4 Flash compare to GPT-4 or Claude?
DeepSeek V4 Flash offers much larger context (1M vs 128k/200k) and output (384k vs ~4k) at a fraction of the cost. However, it is text-only and may have less broad general knowledge or safety tuning.
Does OrcaRouter mark up the price of DeepSeek V4 Flash?
No. OrcaRouter passes through the provider rate with zero markup. You pay $0.14 per 1M input and $0.28 per 1M output exactly as charged by DeepSeek.
How do I call DeepSeek V4 Flash via the OrcaRouter API?
Use the OpenAI-compatible base URL https://api.orcarouter.ai/v1, set the model parameter to "deepseek/deepseek-v4-flash", and include your OrcaRouter API key in the Authorization header.
What data handling policies apply to DeepSeek V4 Flash?
Data passes through OrcaRouter to DeepSeek's servers in China. Review OrcaRouter's privacy policy and DeepSeek's terms. No additional data protections are explicitly offered.
Is DeepSeek V4 Flash multimodal?
No, it only accepts text inputs. For images, audio, or video, you would need to preprocess them into text or use a different model.
What parameters can I set when using DeepSeek V4 Flash?
Standard OpenAI chat completions parameters: model, messages, max_tokens, temperature, top_p, frequency_penalty, presence_penalty, stop, stream, etc. The max_tokens cannot exceed 384,000.
Which use cases are best suited for DeepSeek V4 Flash?
Long-document analysis, code generation with extended reasoning, multi-turn conversations needing deep context, and tasks that produce large outputs such as detailed reports or plans.

Embed this badge

DeepSeek: DeepSeek V4 Flash•$0.22/M in•1941ms p50•via OrcaRouter
HTML <a href="https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash" target="_blank"> <img src="https://www.orcarouter.ai/embed/deepseek/deepseek-v4-flash.svg" alt="DeepSeek: DeepSeek V4 Flash on OrcaRouter" /> </a>
Markdown [![DeepSeek: DeepSeek V4 Flash](https://www.orcarouter.ai/embed/deepseek/deepseek-v4-flash.svg)](https://www.orcarouter.ai/models/deepseek/deepseek-v4-flash)

Model card as data

GET /api/public/models/deepseek/deepseek-v4-flashOpen
Machine-readable:/llms.txt/llms-full.txt