DeepSeek V4 Flash

DeepSeek V4 Flash for low-latency agent, coding and high-volume production workloads

Published
scheduleReleasedJuly 31, 2026

DeepSeek V4 Flash is a DeepSeek V4 model currently backed by DeepSeek-V4-Flash-0731, with a 1M context window, up to 384K output, and thinking and non-thinking modes.

starsCapabilities

codeFunction callingdata_objectStructured output

paymentsContext and pricing

Context limit1,000,000
Max output384,000
Input price¥3/ 1M tokens
Output price¥9/ 1M tokens
Cached input price¥0.1/ 1M tokens

descriptionOverview

Overview

DeepSeek V4 Flash is a DeepSeek V4 model with encyclopedia identifier deepseek-v4-flash, currently backed by DeepSeek-V4-Flash-0731. It supports a 1M context window, up to 384K output, thinking and non-thinking modes, Tool Calls, JSON Output, the Responses API and the Anthropic API.

Versioning

deepseek-v4-flash and deepseek-v4-pro are the official stable API names and track the current versions. Dated entries are encyclopedia snapshots only; DeepSeek does not document them as callable API model parameters. The legacy deepseek-chat and deepseek-reasoner names were retired on July 24, 2026.

lightbulbUse cases

  • Agents and tool use
  • Coding and repository tasks
  • Long-context analysis
  • Chinese Q&A and content generation

thumb_upStrengths

  • 1M context window
  • Up to 384K output
  • Thinking and non-thinking modes
  • OpenAI and Anthropic API compatibility

infoLimitations

  • Stable API names track newer official versions
  • Pinned evaluations should record dated versions
  • Latency and cost require workload testing
  • Official API pricing may increase soon

linkReferences

This content is compiled from official documentation and public sources. Always refer to official documentation for final details