Currently listed through these providers:
Model details
Sarvam-105B
Sarvam-105B is a Mixture-of-Experts model with 106 billion total parameters but only 10.3 billion active parameters engaged per token, achieved through sparse routing that selects 8 experts from a pool of 128 for each forward pass. This architecture lets the model scale to strong reasoning and agentic capabilities while keeping compute costs manageable for real deployment. The attention stack uses MLA-style multi-head attention with decoupled QK projections, rotary position embeddings with an extended theta of 10,000, and SwiGLU activations across 32 layers. A major focus during development was supporting India's linguistic diversity natively—the model achieves state-of-the-art performance across 22 Indian languages built into pretraining itself, enabling authentic multilingual reasoning and generation without relying on translation pipelines.
Sarvam-105B marks a deliberate departure from Sarvam AI's earlier Sarvam-M release, trained from the ground up in India on compute provided under the IndiaAI mission. The full training pipeline included large-scale curated datasets, supervised fine-tuning, and reinforcement learning stages, optimized end-to-end—including tokenization, architecture, execution kernels, and inference systems. The model consistently matches or surpasses major closed-source models across reasoning, programming, and agentic benchmarks, and is already running in production powering the Indus AI assistant for complex reasoning and agentic workflows. Released under Apache 2.0, it provides a viable open-source path for developers seeking competitive frontier-level reasoning with deep Indic language support baked in from the start.
Quick Info
Powered by- Provider
- Sarvam AI
- Model key
- sarvam-105b
- Release date
- Feb 18, 2026
- Last updated
- Mar 6, 2026
- Input modalities
- Output modalities
- Capabilities
Limits
- Output tokens
- 131,072 tokens
- Context window
- 131,072 tokens