Sulat.com
AI models
Vertex logo

Model details

Llama 4 Maverick 17B 128E Instruct

Llama 4 Maverick is built on a mixture-of-experts architecture that activates 17 billion parameters while routing through 128 specialized experts, enabling efficient scaling without proportional compute cost. The model uses early fusion to natively process both text and image inputs together, allowing it to reason about visual content in the same latent space as language. This design prioritizes versatility across chat, knowledge, and code tasks while keeping inference practical for real-world applications.

The instruction-tuned variant has been evaluated against challenging benchmarks including ChartQA (90.0), DocVQA (94.4 anls), and MMLU Pro (59.6), demonstrating strong performance on document understanding and visual reasoning tasks. With a 128K token context window, the model supports applications requiring extended memory, document analysis, and sustained conversation history. The architecture is designed to be highly steerable through system prompts, making it well-suited for building conversational assistants, code generation tools, and enterprise-scale multimodal applications where developers need control over model behavior.

Vertexmeta/llama-4-maverick-17b-128e-instruct-maasllama

Quick Info

Powered by
Provider
Vertex
Model key
meta/llama-4-maverick-17b-128e-instruct-maas
Release date
Apr 29, 2025
Last updated
Apr 29, 2025
Knowledge cutoff
2024-08
AI SDK package
@ai-sdk/openai-compatible
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.35
Output token cost
$1.15

Limits

Output tokens
8,192 tokens
Context window
524,288 tokens

Latest news about Llama 4 Maverick 17B 128E Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Llama 4 Maverick 17B 128E Instruct

More models around Llama 4 Maverick 17B 128E Instruct