Currently listed through these providers:
Model details
DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 continues the Flash-tier naming lineage from the broader DeepSeek family, with the 0731 suffix following the release-date versioning convention familiar from earlier variants. Community discussions on the NVIDIA Developer Forums place the model in two deployment-oriented contexts: a single-node DGX Spark / GB10 thread reporting approximately 1,000 tokens per second of prefill throughput and around 59 tokens per second in a multi-agent serving configuration, and a separate NVIDIA NIM / Models thread indicating the variant is being explored for packaging through NVIDIA's inference distribution platform. These threads are independent third-party benchmarks and packaging discussions rather than an official model card or technical report, so the figures should be read as community-reported reference points rather than vendor-validated performance claims.
As a Flash-tier member of the family, the model targets throughput-sensitive text workloads where lower latency and open-weight deployment flexibility matter more than top-of-range reasoning quality. The combination of community-reported single-node inference speeds on compact GB10 hardware and ongoing NIM packaging activity suggests practical fit for local-agent stacks, multi-agent orchestration scenarios, and cost-aware batch serving where a developer wants to self-host rather than rely on a closed API. Because the available evidence is limited to forum threads rather than official provider documentation, prospective users should treat those throughput numbers as starting points and validate behavior against their own pipelines before committing to production workloads.
Quick Info
Powered by- Provider
- OpenReason
- Model key
- deepseek-ai/deepseek-v4-flash-0731
- Release date
- Jul 31, 2026
- Last updated
- Jul 31, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.1371
- Output token cost
- $0.2743
Limits
- Output tokens
- 384,000 tokens
- Context window
- 1,000,000 tokens
Latest news about DeepSeek V4 Flash 0731
Videos about DeepSeek V4 Flash 0731
More models around DeepSeek V4 Flash 0731
This exact model name is also listed by 38 other providers.