Vercel AI Gateway
You can now access GLM 5V Turbo on Vercel's AI Gateway with no markup and no other provider accounts required.
Model details
GLM 5V Turbo is a multimodal variant in the GLM family designed for agent-driven workflows that need to reason over visual material alongside text, making it well suited for tasks involving design mockups, screenshots, UI recordings, and PDF attachments. A Medium write-up by Agent Native frames it as a multimodal contender, headlining a comparison against Claude Opus 4.6 on multimodal benchmarks shortly after launch, though the supplied excerpts do not include the underlying numbers or methodology. Within Vercel's routing layer, the model is offered alongside its text-only sibling GLM-5-Turbo, with the practical distinction centered on input modality rather than pricing, so teams can pick GLM-5V-Turbo when their pipelines ingest images or documents.
For practical deployment, GLM 5V Turbo reaches developers through Vercel's AI Gateway without markup or extra provider accounts, which simplifies billing and key management for teams already building on that gateway. The model supports reasoning, tool calling, and temperature control, and accepts text, image, and PDF inputs, so it can sit inside agent loops that need to parse UI artifacts or document pages before producing text outputs. The combination of a large context window and multimodal ingest makes it a reasonable choice for agentic applications that mix visual evidence with extended textual reasoning.
Transparent token rates
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Vercel AI Gateway
You can now access GLM 5V Turbo on Vercel's AI Gateway with no markup and no other provider accounts required.
Vercel AI Gateway
An alphaXiv-hosted technical report presents GLM-5V-Turbo as a step toward native foundation models for multimodal agents, arguing that real-world agentic capability requires perceiving, interpreting, and acting over heterogeneous contexts including images, videos, webpages, documents, and GUIs rather than treating vis The report underscores a shift from passive language understanding to active agentic interaction across web browsers, mobile operating systems, and professional software suites, environments the authors describe as inherently multimodal. GLM-5V-Turbo integrates multimodal understanding directly into its reasoning and d