Sulat.com
AI models
Get 10-25% off
Get 10-25% off from Qwen
Alibaba (China) logo

Model details

Qwen2.5-Coder 7B Instruct

Qwen2.5-Coder 7B Instruct belongs to a family of code-specific language models that evolved from the earlier CodeQwen line, built directly on the Qwen2.5 foundation. This architecture employs transformer components with RoPE positional encoding, SwiGLU activation, RMSNorm, and attention QKV bias, organized across 28 layers using group query attention where 28 heads handle queries and 4 shared heads manage key-value projections. The model was designed as a practical code workhorse, focusing on generation, reasoning, and repair tasks while maintaining competence in general language and mathematics. Its 7.61 billion total parameters (6.53 billion excluding embeddings) position it as a mid-range option within a six-model series spanning from 0.5 to 32 billion parameters, offering developers flexibility to match model size to their computational constraints.

The training approach continued pretraining on an extensive corpus exceeding 5.5 trillion tokens encompassing source code, text-code grounding, and synthetic data generated through scalable pipelines with careful data cleaning and balanced mixing. This investment in diverse code-focused training data enabled the series to achieve state-of-the-art performance across more than 10 benchmarks spanning generation, completion, reasoning, and repair tasks, with larger variants in the family demonstrating coding abilities on par with leading proprietary models. The instruction-tuned variant refines this base for interactive use cases including code agents and real-world development workflows, while the model's 131K-token context window supports processing substantial codebases and long documentation. The permissive open-weight licensing reflects a broader strategy to encourage adoption and research contribution across the developer community.

Alibaba (China)qwen2-5-coder-7b-instructqwen

Quick Info

Powered by
Provider
Alibaba (China)
Model key
qwen2-5-coder-7b-instruct
Release date
Nov 1, 2024
Last updated
Nov 1, 2024
Knowledge cutoff
2024-04
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.144
Output token cost
$0.287

Limits

Output tokens
8,192 tokens
Context window
131,072 tokens

Latest news about Qwen2.5-Coder 7B Instruct

No articles yet. Fetch the latest news to show it here.

Videos about Qwen2.5-Coder 7B Instruct

More models around Qwen2.5-Coder 7B Instruct