Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Jun 13, 2026
Input modalities
Output modalities
Capabilities
Recent tweets and retweets from Ollama Cloud
DeepSeek-V4-Flash-0731 is Ollama's fastest growing model ever in token usage. We are scaling capacity in US & Europe.
On Ollama, this model runs with high performance (100tps+) and zero data retention. Your data stays yours.
ollama run deepseek-v4-flash:0731-cloud
Model page: ollama.com/library/deepseek-…
Link
deepseek-v4-flash:0731-cloud
DeepSeek-V4-Flash is a preview of the DeepSeek-V4 series, a Mixture-of-Experts model with 284B total parameters and 13B activated, built for efficient reasoning across a 1M-token context…
DeepSeek-V4-Flash-0731 is now available on Ollama's cloud. This update substantially enhances the model's agentic capabilities:
ollama run deepseek-v4-flash:0731-cloud
Use it with Claude Code:
ollama launch claude --model deepseek-v4-flash:0731-cloud
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.