Currently listed through these providers:
Model details
DeepSeek V4 Flash
DeepSeek V4 Flash sits in the Flash family of compact language models built for fast, efficient deployment. It is published as an open-weight release, letting teams host the weights locally and integrate them into agentic pipelines. The model is available through mainstream inference ecosystems, including NVIDIA's NIM catalog under the deepseek-ai team, which signals broad GPU compatibility and a path to standardized serving. A community thread on the NVIDIA Developer Forums further documents single-DGX Spark running for a date-stamped variant of the model, underscoring its appeal for lightweight local serving rather than large-scale data-center rollouts.
The model is tuned for agent-style workloads: it accepts text input and produces text output, and its capability set includes reasoning, tool calling, temperature control, and structured output, which together support reliable function-calling and JSON-formatted responses. Practically, the model fits workflows where low latency and predictable formatting matter more than the deepest reasoning, such as routing layers, tool orchestration, and high-volume assistants that fan out many short completions. Its lightweight footprint also makes it a practical default when teams want a self-hostable open-weights model that plugs cleanly into existing agent frameworks without the overhead of a flagship-tier system.
Quick Info
Powered by- Provider
- TokenGo
- Model key
- deepseek/deepseek-v4-flash
- Release date
- Apr 24, 2026
- Last updated
- Apr 24, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.098
- Output token cost
- $0.196
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
Latest news about DeepSeek V4 Flash
Videos about DeepSeek V4 Flash
More models around DeepSeek V4 Flash
This exact model name is also listed by 18 other providers.