Currently listed through these providers:
Model details
Llama 4 Maverick 17B FP8
Llama 4 Maverick 17B is built upon a Mixture-of-Experts architecture, utilizing 17 billion parameters to balance performance and efficiency. This design is specifically engineered to excel in high-quality conversational interactions and creative writing workflows. By integrating both text and image understanding, the model serves as a robust foundation for developers looking to build sophisticated AI assistants and interactive applications that require nuanced content generation.
The model benefits from specialized conversational fine-tuning, which sharpens its ability to engage in natural, context-aware dialogue. Its architectural focus on expert-based processing allows it to maintain precision across diverse inputs, making it a strong candidate for applications where visual analysis and textual fluency must work in tandem. As part of the broader Llama 4 family, it provides a scalable solution for developers aiming to implement reliable, multimodal reasoning in their own software environments.
Quick Info
Powered by- Provider
- Deep Infra
- Model key
- meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.80
Limits
- Output tokens
- 16,384 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Llama 4 Maverick 17B FP8 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.