Currently listed through these providers:
Model details
Llama 3.3 70B Instruct
Llama 3.3 70B Instruct is Meta's instruction-tuned 70-billion-parameter text model built for multilingual dialogue, delivered through Vertex AI's Model-as-a-Service channel as an open-weight option. Its underlying architecture follows an optimized auto-regressive transformer design, and the tuned variants are produced through supervised fine-tuning combined with reinforcement learning from human feedback so the responses align with helpfulness and safety preferences. The instruction-tuned text-only release was positioned by Meta to compete favorably with many existing open-source and closed chat models on common industry benchmarks, making it a strong generalist choice for conversational assistants, content generation, and reasoning tasks that do not require image or audio input.
For developers working on long-running documents or extended conversations, the model offers a generous 128,000-token context window that supports substantial document analysis and multi-turn workflows in a single prompt. Practical deployment fit is strongest where teams want a self-hostable or open-weight alternative for production chat applications, multilingual customer support, and retrieval-augmented generation pipelines that benefit from large context spans. Teams planning long-term rollouts should be aware that surrounding hosting providers are beginning to schedule endpoint retirements, so migration paths to newer Llama-family checkpoints or alternative serving platforms are worth evaluating as part of any roadmap tied to this model.
Quick Info
Powered by- Provider
- Vertex
- Model key
- meta/llama-3.3-70b-instruct-maas
- Release date
- Apr 29, 2025
- Last updated
- Apr 29, 2025
- Knowledge cutoff
- 2023-12
- AI SDK package
@ai-sdk/openai-compatible- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.72
- Output token cost
- $0.72
Limits
- Output tokens
- 8,192 tokens
- Context window
- 128,000 tokens
Transparent token rates
Compare llama pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 3.3 70B Instruct
No articles yet. Fetch the latest news to show it here.
Videos about Llama 3.3 70B Instruct
More models around Llama 3.3 70B Instruct
This exact model name is also listed by 24 other providers.