Sulat.com
AI models
Nvidia logo

Model details

Llama-3.2-90B-Vision-Instruct

Llama 3.2 90B Vision Instruct belongs to the Llama family of large language models and is positioned as a multimodal variant designed to jointly process text and images. It accepts both text and image inputs while producing text outputs, making it suitable for vision-grounded conversational tasks such as image captioning, visual question answering, and document or chart interpretation. Third-party aggregator listings confirm the model is publicly accessible through OpenAI-compatible APIs as well as router-style marketplaces, allowing it to be dropped into existing SDK workflows with minimal integration effort.

The model's intended practical strength lies in combining strong language understanding with visual perception, which makes it a natural fit for workflows where users need to ask questions about images, extract structured information from screenshots or scanned documents, or build assistants that reason over mixed text-and-image inputs. Because it is published as a vision-instruct variant, it is best matched to use cases where instruction-following over visual content matters more than raw text-only generation, and where teams prefer a self-hostable, open-weight alternative to closed multimodal APIs for prototyping or production deployment.

Nvidiameta/llama-3.2-90b-vision-instructllama

Quick Info

Powered by
Provider
Nvidia
Model key
meta/llama-3.2-90b-vision-instruct
Release date
Sep 25, 2024
Last updated
Sep 25, 2024
Knowledge cutoff
2023-12
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
8,192 tokens
Context window
128,000 tokens

Latest news about Llama-3.2-90B-Vision-Instruct

Videos about Llama-3.2-90B-Vision-Instruct

Recent tweets and retweets from Nvidia

More models around Llama-3.2-90B-Vision-Instruct