Vercel AI Gateway
BuildFastWithAi's review describes Ling 3.1 Flash as InclusionAI's new hybrid reasoning model released September 29, 2026, a 560B-total-parameter MoE with about 25B active per token, positioned around coding, multi-step analysis, and tool-using agents. The review explicitly distinguishes the announced 1M-token design target from the Vercel AI Gateway-served configuration, which currently exposes 262,144 tokens of context and up to 32,768 output tokens, a difference the author flags as material for production deployment planning. The review consolidates InclusionAI's launch benchmark numbers, including 52.5% on AutomationBench, 68.7% on SkillsBench, 87.9% on CyberGym, 57.9% on Finance Agent v2, 85.5% on DRACO, 40.4% on Terminal-Bench 4, 55.9% on SWE Atlas Codebase QnA, and 65.3% on HealthBench Professional, while noting promotional Vercel access is mentioned through October 13, 2026. It frames the model as a substantial scale jump from Ling 3.0 Flash and warns that the Vercel-served context is materially smaller than InclusionAI's announced design ceiling.