Currently listed through these providers:
Model details
DeepSeek V4 Flash Latest
DeepSeek V4 Flash Latest continues the DeepSeek family of large language models, positioned as a flash-tier variant intended to balance responsiveness with the reasoning behavior seen across recent DeepSeek releases. The model is presented through OpenRouter as a text-to-text chat endpoint that supports structured tool use and configurable sampling, which makes it a fit for assistants, agents, multi-step reasoning pipelines, and retrieval-augmented workflows where a single endpoint can be reused across routing and orchestration layers.</parameter>
Independent distribution via an NVIDIA NIM self-hosted microservice, exposed as nvcr.io/nim/deepseek-ai/deepseek-v4-flash:latest, points to a deployment-ready container image rather than a downloadable public checkpoint, and the same image surfaces an OpenAI-compatible /v1/chat/completions interface on port 8000, suggesting the weights are packaged for inference on NVIDIA hardware stacks. For practitioners, the practical advantage is that V4 Flash Latest can be reached either as a managed API call through OpenRouter or as a local container behind the same request shape, which simplifies client code across hybrid and on-prem environments.</parameter>
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- ~deepseek/deepseek-v4-flash-latest
- Release date
- Aug 1, 2026
- Last updated
- Aug 1, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.05
- Output token cost
- $0.16
Limits
- Output tokens
- 393,216 tokens
- Context window
- 1,310,720 tokens