submodel
DeepSeek-V3-0324 is a 671B parameter Mixture-of-Experts (MoE) model that builds notable updates on top of its predecessor, DeepSeek-V3.
Model details
DeepSeek V3 0324 is a large-scale Mixture-of-Experts language model that pushes further into territory where few open-weight models dare to compete. The parameter count grew from 671 billion to 685 billion, giving the architecture more capacity to handle nuanced reasoning and complex problem-solving. By design, the MoE approach allows the model to activate different expert subnetworks depending on the task, balancing computational efficiency with raw capability. The focus of this update landed squarely on logic reasoning, mathematical problem-solving, and code generation—areas where the sources describe a "quantum leap" rather than an incremental tweak.
This iteration builds directly on the DeepSeek-V3 lineage while targeting the specific weak points identified in earlier versions. Benchmark evaluations back up the claims of improvement: GPQA climbed by over 9 points, AIME jumped nearly 20 points, and LiveCodeBench gained 10 points—numbers that reflect real-world gains in how the model handles graduate-level reasoning and live coding challenges. Function calling accuracy was also sharpened, addressing reliability issues present in the previous V3 release. The model now aligns more closely with the R1 family's writing standards for Chinese-language content, making it suitable for applications that demand both analytical rigor and stylistic quality. For developers and enterprises, the combination of open-source availability, strong reasoning performance, and expanded capabilities positions this model as a practical choice for building AI-powered products that go beyond simple text generation.
submodel
DeepSeek-V3-0324 is a 671B parameter Mixture-of-Experts (MoE) model that builds notable updates on top of its predecessor, DeepSeek-V3.
submodel
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. $0.20 per million input tokens, $0.77 per million output tokens. 163,840 token context window, maximum output of 16,384 tokens. Higher uptime with 6 providers. Includes independent