SiliconFlow
Compare Hunyuan-A13B-Instruct and Ling-flash-2.0 across performance, cost, capabilities, and real-world use cases. See which model fits your needs.
Model details
Ling-flash-2.0 is a Mixture-of-Experts language model built under Ant Group's Ling 2.0 architecture, pairing 100 billion total parameters with a sparse activation strategy that fires only about 6.1 billion parameters per forward pass. This design lets the model punch far above the weight class of typical dense models, delivering reasoning and generation quality comparable to dense models upwards of 30–40 billion parameters while keeping inference more economical. The model is particularly tuned for complex, multi-step tasks: strong performance in mathematical problem-solving, intricate code generation, and frontend development reflects its training on diverse, high-quality data and its engineered expertise routing.
The model was trained on over 20 trillion tokens of curated data and refined through supervised fine-tuning followed by multi-stage reinforcement learning—a pipeline that shapes both factual accuracy and response style. Benchmark results cited from official documentation show competitive scores on multi-disciplinary knowledge reasoning (GPQA-Diamond, MMLU-Pro), advanced math benchmarks (AIME 2025, Omni-MATH), and coding evaluations (LiveCodeBench, CodeForces-Elo), positioning Ling-flash-2.0 as a capable generalist that also excels in technical domains. Released under an MIT license, the model is open-sourced and available in GGUF quantized formats for local deployment, with API access also provided through SiliconFlow for hosted inference.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
SiliconFlow
Compare Hunyuan-A13B-Instruct and Ling-flash-2.0 across performance, cost, capabilities, and real-world use cases. See which model fits your needs.
SiliconFlow
inclusionAI has officially open-sourced Ling-flash-2.0, the third MoE LLM under the Ling 2.0 architecture, following Ling-mini-2.0 and Ring-mini-2.0. The model has 100B total parameters with 6.1B activated (4.8B non-embedding), trained on 20T+ tokens with supervised fine-tuning and multi-stage reinforcement learning. It targets SOTA performance among dense models under 40B parameters while activating only 6B. The Ling 2.0 architecture uses a 1/32 activation-ratio MoE design optimized via Ling Scaling Laws, incorporating expert granularity tuning, shared-expert ratio, attention balance, aux-loss-free plus sigmoid routing, MTP layers, QK-Norm, and Partial-RoPE. Ling-flash-2.0 demonstrates strong complex reasoning on GPQA-Diamond, MMLU-Pro, AIME 2025, Omni-MATH, LiveCodeBench v6, and CodeForces-Elo, with particular strength in code generation and frontend development.