DeepSeek V4 Flash Vision Exp
DeepSeek V4 Flash experimental vision model for image understanding and long context
DeepSeek V4 Flash Vision Exp is an experimental DeepSeek V4 vision model with text and image input, a 1M context window, up to 384K output, and thinking and non-thinking modes.
starsCapabilities
paymentsContext and pricing
descriptionOverview
Overview
DeepSeek V4 Flash Vision Exp is an experimental vision model listed in the official DeepSeek API documentation. Its model ID is deepseek-v4-flash-vision-exp, and its current version is DeepSeek-V4-Flash-Vision-Exp. It supports text and image input, a 1M context window, up to 384K output, Tool Calls, JSON Output, the Responses API and the Anthropic API.
Image billing
Images are converted into tokens based on their dimensions and billed together with text input tokens. Verify actual image token usage against the official vision documentation and API usage response.
lightbulbUse cases
- Image understanding and Q&A
- Screenshot and UI analysis
- Visual document understanding
- Visual information extraction
thumb_upStrengths
- Text and image input
- 1M context window
- Up to 384K output
- Tool Calls and structured output
infoLimitations
- Experimental model whose API and capabilities may change
- Images are converted into billable input tokens
- Vision accuracy requires workload-specific testing
- Pricing and peak rules depend on official documentation
Scan to join WeChat group