Currently listed through these providers:
Model details
Llama 4 Maverick 17B 128E Instruct FP8
Llama 4 Maverick is Meta's multimodal entry in the Llama 4 family, delivered here in an FP8-quantized instruction-tuned variant that retains the visual reasoning capabilities introduced with the broader collection. Microsoft's catalog characterization frames the model as strong at precise image understanding and creative writing, positioning it as a higher-quality, lower-priced alternative to the earlier Llama 3.3 70B. The Hugging Face repository confirms the weights are distributed under the Llama 4 Community License Agreement (Effective Date April 5, 2025), with Meta's official documentation hosted at llama.com/docs/overview, giving teams a clearly licensed, openly shared base to build on rather than a closed proprietary endpoint.
Beyond its text-and-image reasoning profile, the model is designed to fit production multimodal workflows where cost and latency matter more than flagship-scale reasoning. The Foundry listing treats it as a "Direct from Azure" enterprise offering with unified billing and governance, while the open-weights release under Meta's community license lets the same underlying model travel across clouds, including through hosted catalogs that surface attachments, tool calling, and temperature control. In practical terms, Maverick is well suited to vision-grounded assistants, document and chart Q&A, image captioning pipelines, and creative content tasks where pairing tight per-token economics with reliable image comprehension is the deciding factor.
Quick Info
Powered by- Provider
- watsonx.ai
- Model key
- meta-llama/llama-4-maverick-17b-128e-instruct-fp8
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.371
- Output token cost
- $1.484
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Llama 4 Maverick 17B 128E Instruct FP8 pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 4 Maverick 17B 128E Instruct FP8
No articles yet. Fetch the latest news to show it here.
Videos about Llama 4 Maverick 17B 128E Instruct FP8
More models around Llama 4 Maverick 17B 128E Instruct FP8
This exact model name is also listed by 4 other providers.