Currently listed through these providers:
Model details
GLM 5.3 Flash TEE
GLM 5.3 Flash TEE packages Z.AI's natively multimodal architecture—320 billion total parameters with 18 billion active per forward pass—inside a Trusted Execution Environment served by Phala, with Redpill attestation and signed completion receipts appended to every response. This makes the deployment especially attractive for workflows that require cryptographic proof of inference, such as regulated industries, auditable agent pipelines, and enterprise integrations where tamper-evidence matters more than raw throughput. The underlying GLM-5 lineage emphasizes long-context efficiency and precise tool use, so the TEE wrapper preserves the model's intended strengths in coding assistance, agentic task execution, and million-token retrieval-augmented reasoning.
On the public LMArena leaderboard, the model posts an overall score of 1469.4, ranking 38 out of 395 entries and earning category-leading placement in hard prompts at 1495.0 with an Expert tier score of 1514.2, while trailing in creative writing at 1433.3. These figures suggest a model tuned for analytical and instruction-following workloads rather than open-ended generation, aligning well with technical drafting, multi-step problem solving, and structured-output pipelines. The hosted configuration on NanoGPT exposes the full million-token window alongside the complete capability surface, making it a practical choice for teams that need verifiable inference without sacrificing multimodal input or the long context the GLM-5 family was designed to leverage.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- TEE/glm-5.3-flash
- Release date
- Aug 26, 2026
- Last updated
- Aug 26, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.15
- Output token cost
- $0.50
Limits
- Input tokens
- 1,048,576 tokens
- Output tokens
- 131,072 tokens
- Context window
- 1,048,576 tokens