Currently listed through these providers:
Model details
Qwen3 32B TEE
Qwen3 32B TEE is a 32-billion-parameter Qwen3 model packaged for confidential inference, running inside an Intel TDX Trusted Domain so that prompts and responses remain encrypted in memory during processing. This deployment style is aimed at workflows that handle sensitive material such as legal documents, patient records, or financial data, where hardware-isolated execution is a hard requirement. The model is positioned for reasoning, coding, and instruction following, combining a mid-sized parameter footprint with a context window that sources describe as approximately 40K to 41K tokens, making it suitable for long passages of structured text without requiring a much larger mixture-of-experts stack.
In practice, Qwen3 32B TEE is exposed through OpenAI-compatible endpoints, with the gateway listing supporting tool use along with explicit and implicit caching to help control repeated prompt costs. Among TEE-hosted models, it is described as the most widely used option on one confidential-compute platform, reflecting broad adoption for production assistants and agents that need both inference capability and attestation-grade isolation. Its balance of size, instruction-following strength, and confidential execution makes it a practical fit for teams that want a capable general-purpose model but cannot send workloads to standard non-isolated infrastructure.
Quick Info
Powered by- Provider
- Chutes
- Model key
- Qwen/Qwen3-32B-TEE
- Release date
- Apr 1, 2025
- Last updated
- Apr 1, 2025
- Knowledge cutoff
- 2025-04
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.104
- Output token cost
- $0.416
Limits
- Output tokens
- 40,960 tokens
- Context window
- 40,960 tokens