Kilo Gateway
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It...
Model details
DeepSeek V3.1 is a large-scale hybrid model built on a 671B parameter Mixture-of-Experts architecture that activates 37B parameters per forward pass. Designed as a versatile successor to previous iterations, it integrates both thinking and non-thinking modes, allowing users to toggle reasoning behaviors through prompt templates. This architecture is specifically engineered to balance depth and speed, making it a robust choice for demanding applications such as structured tool calling, code generation, and the operation of autonomous search or coding agents.
The model lineage benefits from a two-phase long-context training process and the use of FP8 microscaling to achieve high inference efficiency and stability. By combining the strengths of earlier base models with advanced reasoning capabilities, it provides performance comparable to specialized reasoning models while maintaining faster response times. Its design supports complex, long-horizon workflows, positioning it as a practical tool for researchers and developers who require a balance of high-level reasoning and efficient, large-scale data processing.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Kilo Gateway
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It...