LLM Gateway
ByteMark's Tool Loop benchmark (August 3, 2026) evaluated multi-hop tool calling across 29 tools and 5 requests where no single tool can answer alone. The loop runs for real: the model calls, reads a result, and calls again, with chains locked by identifiers the model cannot know in advance. Cases include dividend yiel Gemini 3.6 Flash appeared in the cost-vs-latency scatter alongside 15 targets including Z.ai GLM-4.7, Gemini 3.5-Lite, Gemini 2.5-Flash, Claude Sonnet 5, Grok 4.5, and DeepSeek v4-pro. The benchmark's design, which punishes loose answers and requires the model to handle arithmetic, geocoding, and identifier resolution
