Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
SiliconFlow (China) logo

Model details

Qwen/Qwen2.5-72B-Instruct

At 72 billion parameters with a full transformer backbone, Qwen2.5-72B-Instruct represents a substantial step in Alibaba's push toward capable instruction-following models. Its architecture incorporates rotary position embeddings (RoPE), SwiGLU activation, and grouped query attention featuring 64 query heads alongside just 8 key-value heads—a design choice that preserves much of multi-head attention's quality while trimming the memory and compute overhead during inference. The 80-layer deep stack and 70 billion non-embedding parameters give the model enough representational capacity to handle complex reasoning chains, extended conversations, and nuanced instruction parsing simultaneously. The design intent centers on strong performance in domains where prior Qwen releases showed room for growth, particularly coding and mathematical problem-solving, while also deepening multilingual proficiency across nearly three dozen languages and sharpening the model's ability to produce cleanly structured JSON and other format-precise outputs that developers increasingly demand.

The model undergoes both pretraining and post-training phases, with evidence pointing to specialized expert models being used to cultivate coding and math capabilities—a deliberate effort to move beyond generalist training alone. Instruction tuning follows this foundation, equipping the model to follow diverse system prompts reliably, maintain coherence across lengthy contexts, and adapt to role-play or condition-setting scenarios that many conversational deployments require. The 128K token context window allows the model to ingest entire codebases, lengthy documents, or multi-turn conversation histories in a single pass, while its 8K+ token generation ceiling supports producing substantial outputs like detailed reports, code implementations, or extended analyses in one go. These attributes make it particularly well-suited for developer-focused applications, data extraction pipelines, and multilingual services that need a model capable of both understanding and generating structured, long-form content with consistency.

SiliconFlow (China)Qwen/Qwen2.5-72B-Instructqwen

Quick Info

Powered by
Provider
SiliconFlow (China)
Model key
Qwen/Qwen2.5-72B-Instruct
Release date
Sep 18, 2024
Last updated
Nov 25, 2025
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.59
Output token cost
$0.59

Limits

Output tokens
4,000 tokens
Context window
33,000 tokens

Transparent token rates

Compare Qwen/Qwen2.5-72B-Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen/Qwen2.5-72B-Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Qwen/Qwen2.5-72B-Instruct

More models around Qwen/Qwen2.5-72B-Instruct