Add corrections, implementation notes, pricing changes, or usage caveats for other readers.
Last updated
Apr 7, 2026
Input modalities
Output modalities
Capabilities
202,800 tokens
Recent tweets and retweets from Baseten
We just launched a new tier of Fast Model APIs, starting with GLM-5.2 Fast.
Our customers run some of the most demanding, real-time workloads. Model APIs are the fastest way to get running with frontier-level models, and our Fast tier has even tighter guardrails around…
Today, we're introducing GLM-5.2 Fast: our GLM-5.2 Model API designed for the most demanding real-time use cases.
The Fast tier delivers 2-3x higher TPS than the standard GLM-5.2 MAPI.
Video
Everyone assumes owning your inference stack is a later-stage problem. these four founders didn't wait.
@QuantumArjun, @thejackobrien, @sadjamz_ , @aquariusacquah, all made that call early. building on it right now with @baseten.
Aug 4th, SF. off the record conversation on…
Discuss this model
Add corrections, implementation notes, pricing changes, or usage caveats for other readers.