Currently listed through these providers:
Model details
Llama-3.2-90B-Vision-Instruct
Llama 3.2 90B Vision Instruct belongs to the Llama family of large language models and is positioned as a multimodal variant designed to jointly process text and images. It accepts both text and image inputs while producing text outputs, making it suitable for vision-grounded conversational tasks such as image captioning, visual question answering, and document or chart interpretation. Third-party aggregator listings confirm the model is publicly accessible through OpenAI-compatible APIs as well as router-style marketplaces, allowing it to be dropped into existing SDK workflows with minimal integration effort.
The model's intended practical strength lies in combining strong language understanding with visual perception, which makes it a natural fit for workflows where users need to ask questions about images, extract structured information from screenshots or scanned documents, or build assistants that reason over mixed text-and-image inputs. Because it is published as a vision-instruct variant, it is best matched to use cases where instruction-following over visual content matters more than raw text-only generation, and where teams prefer a self-hostable, open-weight alternative to closed multimodal APIs for prototyping or production deployment.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- meta/llama-3.2-90b-vision-instruct
- Release date
- Sep 25, 2024
- Last updated
- Sep 25, 2024
- Knowledge cutoff
- 2023-12
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 8,192 tokens
- Context window
- 128,000 tokens