Sulat.com
AI models
IO.NET logo

Model details

Llama 3.2 90B Vision Instruct

The Llama 3.2 90B Vision Instruct is Meta's flagship open-weight vision-language model, built upon the Llama 3.1 70B language backbone with a cross-attention vision adapter that connects to a dedicated vision encoder. This architectural pairing gives the model substantially more capacity for complex reasoning, synthesis, and generation compared to smaller variants. The design intent centers on multi-element visual analysis—tasks that require the model to examine images alongside demanding text generation, such as explaining diagrams, interpreting charts, or conducting detailed scene understanding across extended conversations. The combination of a large language backbone with vision understanding makes it particularly suited for applications where visual context and sophisticated text output must work together seamlessly.

As an instruction-tuned model, Llama 3.2 90B Vision Instruct has been optimized for dialogue use cases that involve visual recognition, image reasoning, captioning, and answering questions about visual content. The fine-tuning process targets the kinds of tasks that power real-world applications, including visual question answering and document-level understanding. Performance evaluations show these instruction-tuned variants outperform many available open-source and closed multimodal models on standard industry benchmarks, reaching comparable results to popular closed models in human assessments of helpfulness and safety. The model is available for commercial use under the Llama 3.2 Community License, making it accessible for developers building visual understanding into agents, productivity tools, or research applications that require sophisticated image comprehension paired with fluent language generation.

IO.NETmeta-llama/Llama-3.2-90B-Vision-Instructllama

Quick Info

Powered by
Provider
IO.NET
Model key
meta-llama/Llama-3.2-90B-Vision-Instruct
Release date
Sep 25, 2024
Last updated
Sep 25, 2024
Knowledge cutoff
2023-12
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.35
Output token cost
$0.40

Limits

Output tokens
4,096 tokens
Context window
16,000 tokens

Latest news about Llama 3.2 90B Vision Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Llama 3.2 90B Vision Instruct

More models around Llama 3.2 90B Vision Instruct