Currently listed through these providers:
Model details
DeepSeek V4 Flash
DeepSeek V4 Flash sits within a wave of open coding models positioned for real agentic software work, and Umans AI includes it in its hosted lineup alongside models like Kimi K3 and GLM 5.2, framing the offering as frontier open models run on infrastructure the provider owns. The platform markets the model as accessible from familiar coding clients and as part of an open stack with no lock-in, emphasizing that open weights can be audited and that prompts and code are not retained on the host side. That positioning suggests the model is intended for teams that want a coding-focused open model served through a managed endpoint rather than self-hosting. The Umans homepage describes the operational challenge of serving recent open architectures, implying the value proposition is reliable inference for demanding coding workloads rather than novel architecture from the host.
An NVIDIA NGC catalog record exists under the nim organization and the deepseek-ai team for a DeepSeek-V4-Flash model, and an NVIDIA Developer Forums thread titled "DeepSeek V4 Flash with Vision" in the DGX Spark / GB10 subforum from August 2026 indicates community discussion of a vision-capable variant running on NVIDIA edge hardware. The forum discussion suggests developers are exploring how a DeepSeek V4 Flash variant can be deployed on compact NVIDIA platforms, which points to interest in lighter, accelerated deployments of the family. For practical fit, the Umans-hosted text model is best understood as a coding-oriented open model suitable for agentic workflows where teams want a swappable hosted endpoint, while the broader DeepSeek V4 Flash family appears to be attracting experimentation on NVIDIA edge devices for local or hybrid setups.
Quick Info
Powered by- Provider
- Umans AI
- Model key
- umans-deepseek-v4-flash-0731
- Release date
- Jul 31, 2026
- Last updated
- Jul 31, 2026
- Knowledge cutoff
- 2025-05
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.28
Limits
- Output tokens
- 393,215 tokens
- Context window
- 1,048,576 tokens
Latest news about DeepSeek V4 Flash
Videos about DeepSeek V4 Flash
More models around DeepSeek V4 Flash
This exact model name is also listed by 18 other providers.