RunInfra
A September 1, 2026 LLM Rumors analysis compares AI inference serving stacks and reports RunInfra's tightly scoped H100 sweep found only a 1.02x throughput difference between vLLM and SGLang at concurrency 256, while Wafer's KernelArena evaluates AI-generated kernels on NVIDIA B200 and AMD MI300X hardware with reported DeepSeek V4 Flash appears in this article only as illustrative commercial pressure rather than as a RunInfra benchmark target: an OpenRouter provider snapshot shows 8.80x spread on input price and 13.21x spread on output price, but the piece cautions that promotions, precision, routing eligibility, and service quality
