Currently listed through these providers:
Model details
Llama 4 Scout 17B 16E Instruct
Llama 4 Scout 17B 16E Instruct is a natively multimodal model built on a Mixture of Experts architecture, utilizing 16 experts to balance performance and efficiency. With 17 billion active parameters out of a 109 billion total, it is designed to handle complex tasks such as summarization, personalization, and detailed reasoning. Its architecture features interleaved Rotary Position Embeddings, which allow the model to maintain coherence across an extended 131.1K token context window. This design makes it particularly effective for analyzing entire codebases, multi-document corpora, and long-form user activity logs in a single inference pass.
The model was developed to excel in both text and image understanding, with its long-context capabilities validated through needle-in-a-haystack retrieval tests and cumulative negative log-likelihood evaluations. By incorporating inference-time temperature scaling of attention, the model achieves robust length generalization. As part of the Llama 4 generation, it serves as a versatile tool for developers needing to integrate precise reasoning and multimodal input processing into their applications. Its combination of a lean active parameter count and high-capacity context window positions it as a strong candidate for tasks requiring deep, context-aware analysis.
Quick Info
Powered by- Provider
- Azure Cognitive Services
- Model key
- llama-4-scout-17b-16e-instruct
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.20
- Output token cost
- $0.78
Limits
- Output tokens
- 8,192 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare Llama 4 Scout 17B 16E Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 4 Scout 17B 16E Instruct
No articles yet. Fetch the latest news to show it here.