Sulat.com
AI models
Cloudflare Workers AI logo

Model details

Llama 4 Scout 17B 16E Instruct

Llama 4 Scout 17B 16E Instruct is an auto-regressive language model built on a mixture-of-experts architecture with 17B activated parameters out of 109B total, routing work across 16 experts to balance compute and capability. It uses early fusion for native multimodality, allowing text and image inputs to be processed together rather than bolted on, which supports visual recognition, image reasoning, captioning, and visual question answering alongside conversational chat, knowledge work, and code generation. Reporting from Groq's documentation puts its knowledge cutoff at August 2024, and the instruction-tuned variant posts benchmark scores of 52.2 on MMLU Pro, 88.8 on ChartQA, and 94.4 ANLS on DocVQA, signaling balanced competence across general reasoning and document understanding tasks.

In practice, the Scout variant is well suited to applications that need to reason across long inputs, including multi-document summarization, personalization from extensive user activity histories, and navigation of large codebases. Third-party cloud listings, such as Microsoft Foundry, frame it as a strong assistant-style model for multilingual commercial and research use, with visual reasoning as a core capability. The combination of a wide context span, native multimodal inputs, and an open-weights posture makes it a flexible foundation for teams building assistants that need to blend document, image, and conversational understanding without committing to a closed proprietary stack.

Cloudflare Workers AI@cf/meta/llama-4-scout-17b-16e-instructllama

Quick Info

Powered by
Provider
Cloudflare Workers AI
Model key
@cf/meta/llama-4-scout-17b-16e-instruct
Release date
Apr 5, 2025
Last updated
Apr 5, 2025
Knowledge cutoff
2024-08
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.27
Output token cost
$0.85

Limits

Output tokens
16,384 tokens
Context window
131,000 tokens

Transparent token rates

Compare Llama 4 Scout 17B 16E Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Llama 4 Scout 17B 16E Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Llama 4 Scout 17B 16E Instruct

More models around Llama 4 Scout 17B 16E Instruct