Skip to content

DeepSeek V4.1 Flash

DeepSeek's open-weights multimodal MoE model with a 1M-token context window.

By DeepSeeksource

Facts read from official pages, last on

Official page

Price

What the gateway charges per million tokens: what you send in, and what the model writes back.

Source: Vercel AI Gateway, checked 2026-09-25.

Input
$0.30per 1M tokens
Output
$1.20per 1M tokensOutput costs 4x the input.

Specs

How much it reads at once, how much it writes back, and how recent its knowledge is.

Sources: Vercel AI Gateway, checked 2026-09-25; Hugging Face, checked 2026-09-25.

Context window
1.05M tokens1,048,576
Max output
33K tokens32,768
Knowledge cutoff
Not listed
Released
Type
Language
Modalities
Takes text, image. Returns text.
License
mit
Gateway ID
deepseek/deepseek-v4.1-flash

Capabilities

What the gateway says this model supports.

Source: Vercel AI Gateway, checked 2026-09-25.

  • Reasoning
  • Tool use
  • Implicit caching
  • Vision
  • Structured output

Data handling

What happens to your prompts, and what you have to switch on.

Source: Vercel AI Gateway, checked 2026-09-25.

ZDRSome providers

Zero data retention

Available on a paid Vercel plan through the AI Gateway, from some providers on it (you turn it on).

NTSome providers

No training on prompts

Offered by some providers on the Vercel AI Gateway, when you turn it on.

Usage

How much people actually run it, as counted where it is served.

downloads
606K606,028 downloads

Counted on Hugging Face over the 30 days to 2026-09-25.

requests
598.9M598,910,903 requests
tokens
41.4T41,413,372,939,368 tokens

Counted on OpenRouter over the 16 days to 2026-09-25.

Also by DeepSeek2

Friday

The brief, by email

Five to seven links an issue, each with what to do about it.

Unsubscribe from any issue.

This page as Markdown