Currently listed through these providers:
Model details
Llama 3.2 90B Vision Instruct
The Llama 3.2 90B Vision Instruct is Meta's flagship open-weight vision-language model, built upon the Llama 3.1 70B language backbone with a cross-attention vision adapter that connects to a dedicated vision encoder. This architectural pairing gives the model substantially more capacity for complex reasoning, synthesis, and generation compared to smaller variants. The design intent centers on multi-element visual analysis—tasks that require the model to examine images alongside demanding text generation, such as explaining diagrams, interpreting charts, or conducting detailed scene understanding across extended conversations. The combination of a large language backbone with vision understanding makes it particularly suited for applications where visual context and sophisticated text output must work together seamlessly.
As an instruction-tuned model, Llama 3.2 90B Vision Instruct has been optimized for dialogue use cases that involve visual recognition, image reasoning, captioning, and answering questions about visual content. The fine-tuning process targets the kinds of tasks that power real-world applications, including visual question answering and document-level understanding. Performance evaluations show these instruction-tuned variants outperform many available open-source and closed multimodal models on standard industry benchmarks, reaching comparable results to popular closed models in human assessments of helpfulness and safety. The model is available for commercial use under the Llama 3.2 Community License, making it accessible for developers building visual understanding into agents, productivity tools, or research applications that require sophisticated image comprehension paired with fluent language generation.
Quick Info
Powered by- Provider
- IO.NET
- Model key
- meta-llama/Llama-3.2-90B-Vision-Instruct
- Release date
- Sep 25, 2024
- Last updated
- Sep 25, 2024
- Knowledge cutoff
- 2023-12
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.35
- Output token cost
- $0.40
Limits
- Output tokens
- 4,096 tokens
- Context window
- 16,000 tokens
Latest news about Llama 3.2 90B Vision Instruct
No articles yet. Fetch the latest news to show it here.