Currently listed through these providers:
Model details
Llama-4-Scout-17B-16E-Instruct-FP8
Llama-4-Scout-17B-16E-Instruct-FP8 belongs to Meta's fourth-generation Llama family, which introduced a mixture-of-experts architecture and positioned Scout as the long-context sibling to the Maverick MoE model. Within that lineup, Scout is designed for tasks that benefit from extended document handling, making it a natural fit for long-form summarization, retrieval-augmented workflows, and code or research assistance where larger context windows matter. As an FP8-quantized, instruction-tuned variant, it is intended to give developers a cost-efficient path to that long-context behavior without sacrificing the instruction-following quality of the Llama 4 release.
The broader Llama series remains Meta's flagship open-weights effort, and Scout inherits that transparency: weights are published and the model is deployable across multiple third-party inference providers. Community discussion around serving MoE architectures with modern quantization formats has been active in inference tooling circles, reflecting the operational maturity needed to host a model of this class. Practically, this listing is well suited for teams already standardizing on a gateway that routes to Meta's open models, who want Scout's long-context strengths in a quantized, instruction-tuned form for production assistants, document Q&A, and multimodal-style pipelines that pair text with attached images.
Quick Info
Powered by- Provider
- Vercel AI Gateway
- Model key
- meta/llama-4-scout
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens
Latest news about Llama-4-Scout-17B-16E-Instruct-FP8
No articles yet. Fetch the latest news to show it here.