DeepSeek
DeepSeek-V4-Flash-Vision-Exp
Specs
- Input price
- $0.14/M
- Output price
- $0.28/M
- Context window
- 1M tokens
DeepSeek's first multimodal model - the Flash speed profile (284B MoE, ~13B active) with native image understanding added under the MIT license.
Capabilities
- Text-plus-image input at Flash-class pricing, MIT-licensed weights
- Vision-aware agents: Chartography 64.3, Agents' Last Exam 27.3
- Text agentic performance held up vs Flash-0731 (Terminal-Bench 83.9)
- DSpark speculative decoding and a 1M-token context window
Best for
- Document-heavy pipelines: scans, screenshots, charts, and photos
- Statistical chart reading for lab trends and utilization dashboards
- High-volume multimodal extraction where closed vision APIs are overkill
Limitations to keep in mind
- Experimental checkpoint - can be superseded quickly on a fast-moving family
- Not a radiology or pathology tool; reads documents, not diagnostics
- Latency under heavy image load not independently benchmarked
HIPAA-compliant hosting
DeepSeek-V4-Flash-Vision-Exp is available under our HIPAA-compliant design-partner program, with a signed Business Associate Agreement (BAA), encryption in transit and at rest, and access controls. We are currently onboarding design partners, with general availability coming soon. Run the model behind a unified OpenRouter-style API and swap to another model with a single parameter.
Pricing in context
Open source models like DeepSeek-V4-Flash-Vision-Exp can be served efficiently on optimized inference infrastructure, with savings passed through to you. Exact savings depend on the model and your volume, but open source inference is typically a fraction of the per-token cost of closed-source frontier models on Azure OpenAI or AWS Bedrock - without cloud egress lock-in or minimum commitments.
Design partner program
Run DeepSeek-V4-Flash-Vision-Exp under HIPAA
Join the waitlist to be prioritized. We'll reach out with a qualification call and early access.