Currently listed through these providers:
Model details
Llama 4 Maverick 17B Instruct
Llama 4 Maverick is built on a Mixture of Experts architecture with 17 billion active parameters distributed across 128 routed experts plus a shared expert, yielding 400 billion total parameters that enable sophisticated reasoning while keeping inference efficient. Meta's early fusion approach treats text and vision tokens together from the start, creating more coherent cross-modal understanding rather than bolting on vision as an afterthought. This positions the model for visual recognition, image reasoning, captioning, coding tasks, and agentic workflows where sustained tool use matters.
Benchmark results show competitive performance against comparable models, with the experimental chat version achieving 1417 on the LMArena leaderboard. The model natively supports synthetic data generation and knowledge distillation, which opens paths for custom fine-tuning without relying on a frozen backbone. With tool-calling capabilities and FP8 quantization, Llama 4 Maverick targets production environments that demand efficient inference without sacrificing multimodality, multilingual reasoning, or the flexibility to power autonomous agents.
Quick Info
Powered by- Provider
- Abacus
- Model key
- meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.14
- Output token cost
- $0.59
Limits
- Output tokens
- 8,192 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare llama pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 4 Maverick 17B Instruct
No articles yet. Fetch the latest news to show it here.
Videos about Llama 4 Maverick 17B Instruct
More models around Llama 4 Maverick 17B Instruct
This exact model name is also listed by 5 other providers.