Requesty
Requesty lists a managed Fireworks AI deployment of Z.ai's GLM-5.2 as `fireworks/glm-5.2-fast`, exposed through its OpenAI-compatible router at `https://router.requesty.ai/v1`. The catalog page documents a 1M-token context window with a 131,072-token max output, chat API type, US serving region, no data retention, no t Provider rates for this Fireworks-hosted GLM-5.2 deployment on Requesty are $2.10 per 1M input tokens and $6.60 per 1M output tokens, with a cached-input rate of $0.21 per 1M and a 3.1x output-to-input ratio. Sample workload figures on the page put 100K input + 10K output at $0.28 and 10M input + 1M output at $27.60 be
