Currently listed through these providers:
Model details
DeepSeek V4 Flash
DeepSeek V4 Flash sits inside the broader DeepSeek Flash family of lightweight language models, designed to balance reasoning quality with very fast inference for production use. The model is text-only on both input and output, and combines reasoning, tool calling, temperature control, and structured output capabilities in a single interface, making it well suited for agentic pipelines, retrieval-augmented generation, and other long-context tasks where dependable tool use matters as much as raw text generation. Its open-weights status lets teams self-host and fine-tune, while the very large context and output budgets available in the API are intended to support extended documents, multi-step workflows, and conversational agents that need to keep a long working memory active in a single call.
In the wider ecosystem, the DeepSeek Flash line has been picked up by NVIDIA's NIM catalog, with an entry for DeepSeek V4 Flash published under the deepseek-ai namespace, signaling that the model is part of the official NIM distribution surface rather than a purely community port. A community benchmark on a single DGX Spark / GB10 station reported a DeepSeek-V4-Flash-0731 variant sustaining roughly 1,000 tokens per second at prefill and around 59 tokens per second under multi-agent serving, which is a useful real-world data point for evaluating local deployment density and concurrency. The zero-listed input and output pricing on Pendra, combined with the open weights, makes the model attractive for high-volume experimentation, internal copilots, and research projects where the long context window and tool-calling behavior can be exercised without immediate cost pressure.
Quick Info
Powered by- Provider
- Pendra
- Model key
- deepseek-v4-flash
- Release date
- Apr 24, 2026
- Last updated
- Apr 24, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
Latest news about DeepSeek V4 Flash
Videos about DeepSeek V4 Flash
More models around DeepSeek V4 Flash
This exact model name is also listed by 18 other providers.