Currently listed through these providers:
Model details
Qwen/Qwen2.5-72B-Instruct
At 72 billion parameters with a full transformer backbone, Qwen2.5-72B-Instruct represents a substantial step in Alibaba's push toward capable instruction-following models. Its architecture incorporates rotary position embeddings (RoPE), SwiGLU activation, and grouped query attention featuring 64 query heads alongside just 8 key-value heads—a design choice that preserves much of multi-head attention's quality while trimming the memory and compute overhead during inference. The 80-layer deep stack and 70 billion non-embedding parameters give the model enough representational capacity to handle complex reasoning chains, extended conversations, and nuanced instruction parsing simultaneously. The design intent centers on strong performance in domains where prior Qwen releases showed room for growth, particularly coding and mathematical problem-solving, while also deepening multilingual proficiency across nearly three dozen languages and sharpening the model's ability to produce cleanly structured JSON and other format-precise outputs that developers increasingly demand.
The model undergoes both pretraining and post-training phases, with evidence pointing to specialized expert models being used to cultivate coding and math capabilities—a deliberate effort to move beyond generalist training alone. Instruction tuning follows this foundation, equipping the model to follow diverse system prompts reliably, maintain coherence across lengthy contexts, and adapt to role-play or condition-setting scenarios that many conversational deployments require. The 128K token context window allows the model to ingest entire codebases, lengthy documents, or multi-turn conversation histories in a single pass, while its 8K+ token generation ceiling supports producing substantial outputs like detailed reports, code implementations, or extended analyses in one go. These attributes make it particularly well-suited for developer-focused applications, data extraction pipelines, and multilingual services that need a model capable of both understanding and generating structured, long-form content with consistency.
Quick Info
Powered by- Provider
- SiliconFlow (China)
- Model key
- Qwen/Qwen2.5-72B-Instruct
- Release date
- Sep 18, 2024
- Last updated
- Nov 25, 2025
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.59
- Output token cost
- $0.59
Limits
- Output tokens
- 4,000 tokens
- Context window
- 33,000 tokens
Transparent token rates
Compare Qwen/Qwen2.5-72B-Instruct pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Qwen/Qwen2.5-72B-Instruct
No articles yet. Fetch the latest news to show it here.