Currently listed through these providers:
Model details
Llama-4-Scout-17B-16E-Instruct-FP8
Scout is designed as a multimodal instruction model for long-running conversations and document-oriented work. Its mixture-of-experts design activates 17 billion parameters per token, combining broad model capacity with selective computation; the sources identify native multimodal handling and a nominal 10-million-token context window, making it suitable for workloads that need to process unusually large sequences in one interaction.
The practical appeal is flexible deployment rather than a single specialized strength. The model supports function calling, JSON-formatted responses, and streaming in the documented gateway setup, while Baseten’s preset illustrates vLLM deployment on four H100 GPUs with a 128,000-token serving configuration. The distinction between the advertised context capacity and a particular serving configuration is important when planning hardware, latency, and production throughput.
Quick Info
Powered by- Provider
- Llama
- Model key
- llama-4-scout-17b-16e-instruct-fp8
- Release date
- Apr 5, 2025
- Last updated
- Apr 5, 2025
- Knowledge cutoff
- 2024-08
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 4,096 tokens
- Context window
- 128,000 tokens
Latest news about Llama-4-Scout-17B-16E-Instruct-FP8
No articles yet. Fetch the latest news to show it here.