Currently listed through these providers:
Model details
Llama 4 Scout 17B 16E Instruct
Llama 4 Scout 17B 16E Instruct is an auto-regressive language model built on a mixture-of-experts architecture with 17B activated parameters out of 109B total, routing work across 16 experts to balance compute and capability. It uses early fusion for native multimodality, allowing text and image inputs to be processed together rather than bolted on, which supports visual recognition, image reasoning, captioning, and visual question answering alongside conversational chat, knowledge work, and code generation. Reporting from Groq's documentation puts its knowledge cutoff at August 2024, and the instruction-tuned variant posts benchmark scores of 52.2 on MMLU Pro, 88.8 on ChartQA, and 94.4 ANLS on DocVQA, signaling balanced competence across general reasoning and document understanding tasks.
In practice, the Scout variant is well suited to applications that need to reason across long inputs, including multi-document summarization, personalization from extensive user activity histories, and navigation of large codebases. Third-party cloud listings, such as Microsoft Foundry, frame it as a strong assistant-style model for multilingual commercial and research use, with visual reasoning as a core capability. The combination of a wide context span, native multimodal inputs, and an open-weights posture makes it a flexible foundation for teams building assistants that need to blend document, image, and conversational understanding without committing to a closed proprietary stack.
Quick Info
Powered by- Provider
- Cloudflare Workers AI
- Model key
- @cf/meta/llama-4-scout-17b-16e-instruct
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.27
- Output token cost
- $0.85
Limits
- Output tokens
- 16,384 tokens
- Context window
- 131,000 tokens
Transparent token rates
Compare Llama 4 Scout 17B 16E Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 4 Scout 17B 16E Instruct
No articles yet. Fetch the latest news to show it here.
Videos about Llama 4 Scout 17B 16E Instruct
More models around Llama 4 Scout 17B 16E Instruct
This exact model name is also listed by 2 other providers.