submodel
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It...
Model details
DeepSeek V3.1 is a large hybrid reasoning model built around a flexible dual-mode design that lets users toggle between deep thinking and rapid response modes through simple template switching. With 671 billion parameters and 37 billion activated during inference, it extends its predecessor's foundation through a two-phase long-context training process that pushes effective context out to 128,000 tokens. The architecture leverages FP8 microscaling to keep inference efficient at scale. What sets this model apart is its ability to handle both analytical complexity and everyday tasks within a single unified system, making it suited for research workflows, software engineering, and autonomous agent tasks that demand sustained reasoning alongside quick turnarounds.
The V3.1 iteration builds on the V3 base through continued pretraining on 840 billion additional tokens focused specifically on long-context extension, followed by targeted post-training that sharpens tool use and multi-step agent capabilities. This refinement yields performance on difficult benchmarks comparable to DeepSeek-R1 while responding more quickly, particularly when the thinking mode is engaged. Enhanced code agent and search agent performance shows up in improved results on SWE-Bench and Terminal-Bench, reflecting genuine gains in practical tool-driven tasks rather than just benchmark scores. The model's support for structured tool calling and the Anthropic API format makes it straightforward to integrate into production pipelines, positioning it as a practical upgrade for developers building complex automation systems.
submodel
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It...
This exact model name is also listed by 16 other providers.