Currently listed through these providers:
Model details
Meta Llama 4 Scout 17B 16E
Meta Llama 4 Scout represents a natively multimodal approach within the Llama family, built on an auto-regressive transformer core enhanced by a 16-expert Mixture-of-Experts design. With 17 billion active parameters routing through specialized experts, the architecture allows the model to maintain strong capability while optimizing for inference efficiency rather than activating the entire model for every token. This efficiency-first design makes it particularly well-suited for applications where both performance and resource constraints matter. The model was developed in collaboration with AI Kosh, Meta, and Sarvam, bringing together expertise for instruction-tuned interactions, advanced natural language generation, and multimodal reasoning tasks including image understanding, captioning, and visual question answering.
The instruction-tuning process shapes the model for assistant-style interactions while retaining its native multimodal strengths. Its broad language support spans English, Hindi, French, German, Spanish, Arabic, Indonesian, Italian, Portuguese, and Tagalog, reflecting a multilingual focus that extends the model's practical reach. The expert routing mechanism underpins competitive performance across text and image understanding, enabling the model to handle coding tasks, tool-calling scenarios, and agentic workflows. The combination of efficient MoE processing with multimodal reasoning positions this model as a practical choice for developers building interactive, multilingual AI applications.
Quick Info
Powered by- Provider
- Helicone
- Model key
- llama-4-scout
- Release date
- Jan 1, 2025
- Last updated
- Jan 1, 2025
- Knowledge cutoff
- 2025-01
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.08
- Output token cost
- $0.30
Limits
- Output tokens
- 8,192 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Meta Llama 4 Scout 17B 16E pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Meta Llama 4 Scout 17B 16E
No articles yet. Fetch the latest news to show it here.