Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
OpenRouter logo

Model details

Llama 4 Maverick

Llama 4 Maverick is a natively multimodal large language model built on a sparse mixture-of-experts architecture that routes requests across 128 specialized experts while activating only 17 billion parameters per forward pass, keeping computational costs manageable despite the model's 400 billion total parameters. The architecture uses grouped-query attention with 96 query heads and 8 key-value heads across 120 transformer layers, paired with Swish activation and RMS normalization for stable training and inference. Its early fusion design enables seamless integration of text and image inputs, supporting a one-million-token context window that accommodates long documents, multi-image conversations, or extended reasoning chains. This design positions Maverick for demanding applications like document understanding, visual question answering, and complex code repositories where broad context and multimodal reasoning are essential.

The model was instruction-tuned to behave like a capable assistant, with explicit optimization for image reasoning, multilingual interaction across twelve languages, and code generation tasks. Training drew from a curated blend of public datasets, licensed corpora, and Meta's own platform data spanning roughly 22 trillion tokens, culminating in a knowledge cutoff in mid-2024. Meta employed its GOAT (Generative Offensive Agent Testing) framework during development to simulate adversarial scenarios and strengthen alignment through automated red teaming. Safety tools like Prompt Guard and Llama Guard were open-sourced alongside the model. Released under the Llama 4 Community License, Maverick delivers competitive performance against closed models on benchmarks covering coding, visual reasoning, and general language understanding, making it a strong choice for developers and researchers who need advanced multimodal capabilities with the flexibility of open weights.

OpenRoutermeta-llama/llama-4-maverickllama

Quick Info

Powered by
Provider
OpenRouter
Model key
meta-llama/llama-4-maverick
Release date
Apr 5, 2025
Last updated
Apr 5, 2025
Knowledge cutoff
2024-08-31
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.1875
Output token cost
$0.6525

Limits

Output tokens
16,384 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Llama 4 Maverick pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Llama 4 Maverick

OpenRouter

CoverageBenchmark

The PickModel directory entry for Llama 4 Maverick tags it as a frontier-tier, multimodal, open-source, mixture-of-experts model from Meta with a 1,048,576-token context window, knowledge cutoff of 2024-08, and a release date of 2025-04. It lists vision (image input) support, function/tool calling, JSON mode, streaming Reported pricing is $0.15 per 1M input tokens and $0.60 per 1M output tokens as list prices, which may vary by provider plan. Listed benchmarks include MMLU 88.2%, GPQA Diamond 58.3%, MATH 79.1%, and SWE-bench Verified 32.1%, with AIME 2025, GSM8K, HumanEval, LiveCodeBench, and HellaSwag marked as not yet published. Th

OpenRouter

Official sourceComparison

OpenRouter's official comparison page pits meta-llama/llama-4-maverick against meta-llama/llama-4-scout and lets developers switch between the two via a model slug change rather than a new integration. The page specifies a context window of 1,048,576 tokens for Llama 4 Maverick and 10,000,000 tokens for Llama 4 Scout, The page also surfaces OpenRouter's flagship model curation — Claude Fable 5, Gemini 3.1 Pro Preview, and GPT-5.5 — alongside affordability, code, and reasoning tracks (DeepSeek V4 Flash 0423, Gemini 2.5 Flash Lite, Hy3 preview, Claude Opus 4.7, Gemini 3 Flash Preview, Kimi K2.6, R1 0528, Gemini 2.5 Pro, o4 Mini). Pric

Videos about Llama 4 Maverick

More models around Llama 4 Maverick