Currently listed through these providers:
Model details
Llama 4 Maverick 17B 128E Instruct
Llama 4 Maverick is built on a mixture-of-experts architecture that activates 17 billion parameters while routing through 128 specialized experts, enabling efficient scaling without proportional compute cost. The model uses early fusion to natively process both text and image inputs together, allowing it to reason about visual content in the same latent space as language. This design prioritizes versatility across chat, knowledge, and code tasks while keeping inference practical for real-world applications.
The instruction-tuned variant has been evaluated against challenging benchmarks including ChartQA (90.0), DocVQA (94.4 anls), and MMLU Pro (59.6), demonstrating strong performance on document understanding and visual reasoning tasks. With a 128K token context window, the model supports applications requiring extended memory, document analysis, and sustained conversation history. The architecture is designed to be highly steerable through system prompts, making it well-suited for building conversational assistants, code generation tools, and enterprise-scale multimodal applications where developers need control over model behavior.
Quick Info
Powered by- Provider
- Vertex
- Model key
- meta/llama-4-maverick-17b-128e-instruct-maas
- Release date
- Apr 29, 2025
- Last updated
- Apr 29, 2025
- Knowledge cutoff
- 2024-08
- AI SDK package
@ai-sdk/openai-compatible- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.35
- Output token cost
- $1.15
Limits
- Output tokens
- 8,192 tokens
- Context window
- 524,288 tokens
Latest news about Llama 4 Maverick 17B 128E Instruct
No articles yet. Fetch the latest news to show it here.