Currently listed through:
Model details
Magistral Small
Magistral Small belongs to the Magistral family of reasoning-focused models, designed as a compact option that prioritizes transparent chain-of-thought inference over raw scale. Rather than training a model from scratch, the Magistral approach adds reasoning capabilities on top of an existing base—applying supervised fine-tuning from Magistral Medium traces and reinforcement learning to teach the model to produce long, parseable thinking chunks before arriving at an answer. This pipeline produces a 24-billion-parameter model that can run locally on a single RTX 4090 graphics card or a 32GB RAM MacBook once quantized, making it practical for developers who need reasoning behavior without depending on large cloud deployments.
In third-party agent evaluations, Magistral Small 2506 earned a place on the Galileo Agent Leaderboard at rank 22, where its standout strengths are response speed and cost efficiency: near-instant replies at roughly $0.03 per session. Independent commentary on the same leaderboard, however, flags weaker agent-side behavior, noting that action completion and tool selection accuracy lag behind its latency advantages. The trade-off positions Magistral Small as a strong fit for latency-sensitive, single-pass reasoning workloads where local deployment and fast answers matter more than robust multi-step agent orchestration.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- mistralai/Magistral-Small-2506
- Release date
- Jun 10, 2025
- Last updated
- Jun 10, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.50
- Output token cost
- $1.50
Limits
- Output tokens
- 64,000 tokens
- Context window
- 128,000 tokens