Sulat.com
AI models
Nebius Token Factory logo

Model details

DeepSeek V4 Flash 0731

DeepSeek V4 Flash 0731 is an autoregressive Mixture-of-Experts Transformer built for advanced reasoning and agentic work. Its architecture combines Compressed Sparse Attention and Heavily Compressed Attention with Manifold-Constrained Hyper-Connections, while placing DeepSeek’s DSpark speculative decoding module in the same checkpoint. The MIT-licensed release checkpoint is designed to support demanding text tasks, including tool-using assistants, software engineering, and complex problem solving.

The model’s practical profile centers on reasoning-intensive agents and tool use rather than multimodal work. Published text-agent evaluations cover terminal tasks, repository generation, software engineering, tool interaction, and automation, giving useful evidence for workflows that need multi-step execution. Its sparse hybrid attention and bundled speculative decoding make it a strong fit for applications that value capable reasoning alongside efficient text generation.

Nebius Token Factorydeepseek-ai/DeepSeek-V4-Flash-0731deepseek-flash

Quick Info

Powered by
Provider
Nebius Token Factory
Model key
deepseek-ai/DeepSeek-V4-Flash-0731
Release date
Jul 31, 2026
Last updated
Jul 31, 2026
Knowledge cutoff
2025-05
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.14
Output token cost
$0.28

Limits

Output tokens
1,024,000 tokens
Context window
1,024,000 tokens

Transparent token rates

Compare DeepSeek V4 Flash 0731 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about DeepSeek V4 Flash 0731

Nebius Token Factory

CoverageBenchmark

InferenceBench's Nebius Token Factory dashboard surfaces recent 24-hour performance deltas for DeepSeek V4 Flash served on Nebius, including a throughput ↓58% reading on one window and a reliability recovery 96.8% → 100% on another. The same dashboard shows positive movement on related models routed via Nebius, includi Comparison readings also place DeepSeek V4 Flash on other providers in the same window: Scaleway throughput ↑359% with TTFT ↓39%, Fireworks AI throughput ↓16% with TTFT ↑17%, and Cortecs throughput ↑53%, alongside an OpenRouter reliability recovery 96.8% → 100%. For developers routing DeepSeek V4 Flash on Nebius Token

Nebius Token Factory

CoverageAnalysis

Artificial Analysis published a dedicated evaluation of DeepSeek V4 Flash 0731 on 2026-07-31, reporting an Intelligence Index score of 50 — a 10-point jump over the April 2026 DeepSeek V4 Flash (40) and one point behind GPT-5.6 Luna (max, 51). The model retains the same architecture and pricing as its predecessor, with The biggest gains are agentic: GDPval-AA v2 Elo rises from 1189 to 1559 (second-highest open-weights score behind Kimi K3's 1687 and ahead of GLM-5.2's 1510), Terminal-Bench 2.1 climbs 17 points to 79%, and τ³-Bench Banking improves 8 points to 31%. AA-Omniscience improves from -23 to -16 purely through reduced halluci

RunInfra

CoverageBenchmark

OpenRouter's page for deepseek/deepseek-v4-flash-0731 confirms the July 31, 2026 release of DeepSeek V4 Flash 0731 as a sparse mixture-of-experts model with 284B total parameters and 13B active, a 1M token context window, and a stated base price of $0.05 input / $0.16 output per 1M tokens with text in/out modalities. D As a router, OpenRouter enumerates many downstream inference hosts (OpenInference, Relace, DigitalOcean, DeepInfra, Together, Baseten, CoreWeave, Fireworks, Cloudflare, DeepSeek first-party, and others) with per-host latency, throughput, uptime, and pricing tiers for serving DeepSeek V4 Flash 0731. These figures reflec

Nebius Token Factory

Coverage

DeepSeek's official API changelog documents two V4-family updates that frame the environment around V4 Flash 0731: on 2026-08-13 the GA release of DeepSeek-V4-Pro rolled out with materially enhanced agent capabilities (Terminal Bench 2.1 87.9, NL2Repo 61.5, Cybergym 83.3, Toolathlon-Verified 74.1, DSBench-FullStack 71. On 2026-08-21, DeepSeek shipped DeepSeek-V4-Flash-Vision-Exp, an experimental multimodal vision understanding variant accessed via model='deepseek-v4-flash-vision-exp'. Its reported scores include Terminal Bench 2.1 83.9, NL2Repo 57.7, DeepSWE 59.3, DSBench-Hard 63.6, ApexBench Pass@1 36.5, Agents' Last Exam 27.3, Char

Videos about DeepSeek V4 Flash 0731

More models around DeepSeek V4 Flash 0731