Currently listed through these providers:
Model details
Llama 4 Scout 17B 16E Instruct
Llama 4 Scout 17B 16E Instruct is a Mixture of Experts model with 17 billion active parameters distributed across 16 experts, totaling 109 billion parameters—substantially leaner than its Maverick sibling. Built natively multimodal from the ground up, it processes text, images, and video frames in a unified architecture rather than bolting on vision capabilities as an afterthought. Its defining architectural leap is an extended context window supporting single-pass processing of entire codebases, multi-document corpora, or months of user activity logs. The design uses interleaved Rotary Position Embeddings (iRoPE), where most layers apply standard positional encoding while others interleave attention layers without positional embeddings, enhanced by inference-time temperature scaling to generalize effectively across its full context length.
The model was validated through needle-in-a-haystack retrieval tests and cumulative negative log-likelihood evaluations spanning its full 131.1K token context—roughly 7.5 million words or about 25 novels in a single call. Its sweet spot lies in multi-document summarization, parsing extensive activity logs for personalized task reasoning, and conducting deep analysis across vast code repositories. For vision tasks, it handles visual recognition, image reasoning, and captioning alongside natural language understanding. The model supports twelve languages including Arabic, English, French, German, Hindi, Portuguese, and Spanish, making it viable for international applications. Released as a founding member of the Llama 4 generation, it targets developers and teams needing long-context understanding paired with multimodal comprehension at a scale that remains practical to deploy.
Quick Info
Powered by- Provider
- Azure
- Model key
- llama-4-scout-17b-16e-instruct
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.78
Limits
- Output tokens
- 8,192 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare Llama 4 Scout 17B 16E Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 4 Scout 17B 16E Instruct
No articles yet. Fetch the latest news to show it here.
Videos about Llama 4 Scout 17B 16E Instruct
More models around Llama 4 Scout 17B 16E Instruct
This exact model name is also listed by 2 other providers.