Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba logo

Model details

Qwen2.5 72B Instruct

Qwen2.5 72B Instruct is a large causal language model built on a transformer foundation enhanced with RoPE positional encoding, SwiGLU activations, RMSNorm, and attention QKV bias. The architecture spans 80 layers with 72.7 billion total parameters (70 billion non-embedding), leveraging grouped query attention with 64 query heads and 8 key-value heads to balance capability and efficiency. This model was designed as part of a broader effort to bring substantial improvements in coding, mathematics, and general knowledge compared to its predecessor, drawing on specialized expert model training in these domains. Its ability to handle structured inputs like tables, generate structured outputs especially JSON, and maintain coherence across very long documents reflects both architectural choices and targeted post-training refinements.

The post-training pipeline combines pretraining with instruction-tuning, building on learnings from earlier Qwen generations to strengthen how the model follows complex instructions and maintains consistency under varied system prompts. Multilingual support extending across 29 languages allows deployment in diverse global contexts, while improved resilience to prompt diversity makes it suitable for role-play applications, conditional chatbot scenarios, and other use cases requiring reliable instruction adherence. Performance benchmarks show this model scoring above average on intelligence metrics among comparable open-weight models, with particular strength demonstrated in code generation tasks like LiveCodeBench. The combination of a massive context window, strong instruction fidelity, and multilingual versatility positions it as a practical choice for developers building complex interactive applications, data processing pipelines, or multilingual conversational systems.

Alibabaqwen2-5-72b-instructqwen

Quick Info

Powered by
Provider
Alibaba
Model key
qwen2-5-72b-instruct
Release date
Sep 1, 2024
Last updated
Sep 1, 2024
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.40
Output token cost
$5.60

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Transparent token rates

Compare Qwen2.5 72B Instruct pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Qwen2.5 72B Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Qwen2.5 72B Instruct

More models around Qwen2.5 72B Instruct