DeepSeek V4.1 Flash
DeepSeek's open-weights multimodal MoE model with a 1M-token context window.
Price
What the gateway charges per million tokens: what you send in, and what the model writes back.
Source: Vercel AI Gateway, checked 2026-09-25.
- Input
- $0.30per 1M tokens
- Output
- $1.20per 1M tokensOutput costs 4x the input.
Specs
How much it reads at once, how much it writes back, and how recent its knowledge is.
Sources: Vercel AI Gateway, checked 2026-09-25; Hugging Face, checked 2026-09-25.
- Context window
- 1.05M tokens1,048,576
- Max output
- 33K tokens32,768
- Knowledge cutoff
- Not listed
- Released
- Type
- Language
- Modalities
- Takes text, image. Returns text.
- License
- mit
- Gateway ID
deepseek/deepseek-v4.1-flash
Capabilities
What the gateway says this model supports.
Source: Vercel AI Gateway, checked 2026-09-25.
- Reasoning
- Tool use
- Implicit caching
- Vision
- Structured output
Data handling
What happens to your prompts, and what you have to switch on.
Source: Vercel AI Gateway, checked 2026-09-25.
ZDRSome providers
Zero data retention
Available on a paid Vercel plan through the AI Gateway, from some providers on it (you turn it on).
NTSome providers
No training on prompts
Offered by some providers on the Vercel AI Gateway, when you turn it on.
Usage
How much people actually run it, as counted where it is served.
- downloads
- 606K606,028 downloads
Counted on Hugging Face over the 30 days to 2026-09-25.
- requests
- 598.9M598,910,903 requests
- tokens
- 41.4T41,413,372,939,368 tokens
Counted on OpenRouter over the 16 days to 2026-09-25.
Also by DeepSeek
Friday
The brief, by email
Five to seven links an issue, each with what to do about it.
Unsubscribe from any issue.