NovitaAI
A third-party guide on DeepSeek-Prover-V2-671B confirms the model was released on April 30, 2025, as the next-generation automated theorem proving model in DeepSeek's open-weight lineup. It is built on the same 671 billion-parameter Mixture-of-Experts (MoE) architecture that powers DeepSeek-V3, with an estimated 37 bil Key reported specifications include a context length of approximately 128,000 tokens to accommodate lengthy proofs and complex reasoning chains, and a likely Multi-Head Latent Attention (MLA) mechanism inherited from DeepSeek-V2 that compresses the KV cache to reduce RAM and VRAM requirements. The guide notes the model