LLM Gateway
InclusionAI released Ling-3.0-flash-VL in September 2026 as an open-weights multimodal model under an MIT license, supporting text, image, and video inputs with text output. It uses a Mixture-of-Experts architecture with 124B total parameters and 5.5B active parameters per token during inference. The page reports a 262k token context window and reasoning capability, with weights available on Hugging Face. On the Artificial Analysis Intelligence Index the model scores 25, placing it well above the 8 median for comparable open-weights models in its size class, while generating a somewhat verbose 160M tokens. Throughput is 142.6 output tokens per second, faster than the 133 median. Pricing is reported at $0.075 per 1M input tokens and $0.22 per 1M output tokens, with an 80% cache discount.