Llama 4 Maverick is a natively multimodal large language model built on a sparse mixture-of-experts architecture that routes requests across 128 specialized experts while activating only 17 billion parameters per forward pass, keeping computational costs manageable despite the model's 400 billion total parameters. The architecture uses grouped-query attention with 96 query heads and 8 key-value heads across 120 transformer layers, paired with Swish activation and RMS normalization for stable training and inference. Its early fusion design enables seamless integration of text and image inputs, supporting a one-million-token context window that accommodates long documents, multi-image conversations, or extended reasoning chains. This design positions Maverick for demanding applications like document understanding, visual question answering, and complex code repositories where broad context and multimodal reasoning are essential.
The model was instruction-tuned to behave like a capable assistant, with explicit optimization for image reasoning, multilingual interaction across twelve languages, and code generation tasks. Training drew from a curated blend of public datasets, licensed corpora, and Meta's own platform data spanning roughly 22 trillion tokens, culminating in a knowledge cutoff in mid-2024. Meta employed its GOAT (Generative Offensive Agent Testing) framework during development to simulate adversarial scenarios and strengthen alignment through automated red teaming. Safety tools like Prompt Guard and Llama Guard were open-sourced alongside the model. Released under the Llama 4 Community License, Maverick delivers competitive performance against closed models on benchmarks covering coding, visual reasoning, and general language understanding, making it a strong choice for developers and researchers who need advanced multimodal capabilities with the flexibility of open weights.