Currently listed through these providers:
Model details
Llama 4 Maverick 17B 128E Instruct
Llama 4 Maverick 17B 128E Instruct is Meta's mixture-of-experts flagship in the Llama 4 line, built around an architecture that activates 17 billion parameters from a much larger 128-expert pool per token. This expert-routing design lets the model deliver frontier-class quality while keeping per-token compute closer to a mid-sized model. It is part of Meta's push toward natively multimodal open-weight models, accepting both text and image inputs and producing text along with structured tool-call outputs, which makes it well suited to assistant, agentic, and document-understanding workflows where visual context matters.
In practical use, Maverick is positioned as a versatile mid-to-frontier chat model rather than a narrow specialist. Its open-weight availability allows researchers and builders to fine-tune, distill, or self-host the model for domain-specific assistants, while its tool-calling and structured-output strengths make it a good fit for orchestrating multi-step agents, retrieval pipelines, and code or API-heavy automation. Compared with denser Llama variants, the MoE layout offers a better quality-versus-cost trade-off for long, mixed-modality conversations, and Meta has signaled ongoing investment in the Llama 4 family, so further post-trained or instruction-refined checkpoints are likely to expand its range of supported tasks over time.
Quick Info
Powered by- Provider
- IO.NET
- Model key
- meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
- Release date
- Jan 15, 2025
- Last updated
- Jan 15, 2025
- Knowledge cutoff
- 2024-12
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.60
Limits
- Output tokens
- 4,096 tokens
- Context window
- 430,000 tokens
Latest news about Llama 4 Maverick 17B 128E Instruct
No articles yet. Fetch the latest news to show it here.