Currently listed through these providers:
Model details
Llama 4 Scout 17B Instruct
Meta's Llama 4 Scout 17B Instruct is built on a mixture-of-experts architecture that activates only 17 billion of its 109 billion total parameters during each forward pass, using 16 routed experts to deliver efficient inference without sacrificing capability. This design philosophy prioritizes high computational efficiency, making the model practical for local deployment and commercial applications where resource constraints matter. Native multimodal input support, powered by an early fusion approach that weaves text and image tokens together from the ground up, enables seamless visual reasoning and image understanding rather than bolted-on vision add-ons. The model is instruction-tuned specifically for assistant-style interaction, supporting multilingual output across twelve languages for both text and code generation.
The training corpus for Llama 4 Scout encompasses approximately 40 trillion tokens, with the model's knowledge extending through August 2024, giving it a substantial foundation for general reasoning and domain knowledge. Instruction tuning shapes the base model into a specialized assistant capable of multilingual chat, captioning, visual reasoning, and code generation across multiple programming languages. The open-weights approach under the Llama 4 Community License invites developers to deploy, fine-tune, and build upon the model for custom applications. Its exceptional context length reaching into the millions of tokens positions it well for long-document tasks, research workflows, and scenarios where processing extensive conversational history or large codebases becomes essential.
Quick Info
Powered by- Provider
- DevPass (LLM Gateway)
- Model key
- llama-4-scout-17b-instruct
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.18
- Output token cost
- $0.59
Limits
- Output tokens
- 2,048 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare llama pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 4 Scout 17B Instruct
No articles yet. Fetch the latest news to show it here.