Currently listed through these providers:
Model details
Llama 3.2 3B Instruct
Llama 3.2 3B Instruct is a small, instruction-tuned text model from Meta, part of the Llama 3 family and aimed at scalable assistant and agentic use cases. It is distributed under the Llama 3.2 Community License as an open-weights release, with multiple GGUF quantizations (Q4_K_M, Q5_K_M, Q6_K, and Q8_0) and an Ollama distribution, making it straightforward to run in constrained local environments. The model is described as multilingual and text-only, building on the prior Llama lineage and positioned as a lightweight foundation model for developers who need efficient on-device deployment rather than a large frontier system.
Practically, the 3B Instruct variant fits workflows that need a responsive instruction-following model with a small footprint, such as lightweight chatbots, prototyping, summarization, and tool-augmented agents running on consumer hardware. The range of available GGUF quantizations lets users trade quality against memory, and the Ollama packaging simplifies local setup. As a Meta-released model in the Llama 3 series, it benefits from the broader ecosystem of community fine-tunes and integrations, while staying focused on text generation rather than multimodal tasks.
Quick Info
Powered by- Provider
- Pioneer
- Model key
- meta-llama/Llama-3.2-3B-Instruct
- Release date
- Aug 31, 2024
- Last updated
- Sep 25, 2024
- Knowledge cutoff
- 2023-12-31
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.335
Limits
- Output tokens
- 80,000 tokens
- Context window
- 131,072 tokens
Transparent token rates
Compare Llama 3.2 3B Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Llama 3.2 3B Instruct
No articles yet. Fetch the latest news to show it here.