Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Deep Infra logo

Model details

Llama 4 Maverick 17B FP8

Llama 4 Maverick 17B is built upon a Mixture-of-Experts architecture, utilizing 17 billion parameters to balance performance and efficiency. This design is specifically engineered to excel in high-quality conversational interactions and creative writing workflows. By integrating both text and image understanding, the model serves as a robust foundation for developers looking to build sophisticated AI assistants and interactive applications that require nuanced content generation.

The model benefits from specialized conversational fine-tuning, which sharpens its ability to engage in natural, context-aware dialogue. Its architectural focus on expert-based processing allows it to maintain precision across diverse inputs, making it a strong candidate for applications where visual analysis and textual fluency must work in tandem. As part of the broader Llama 4 family, it provides a scalable solution for developers aiming to implement reliable, multimodal reasoning in their own software environments.

Deep Inframeta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8llama

Quick Info

Powered by
Provider
Deep Infra
Model key
meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
Release date
Apr 5, 2025
Last updated
Apr 5, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.20
Output token cost
$0.80

Limits

Output tokens
16,384 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Llama 4 Maverick 17B FP8 pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Llama 4 Maverick 17B FP8

Videos about Llama 4 Maverick 17B FP8

More models around Llama 4 Maverick 17B FP8