Trinity Large Thinking is a sparse Mixture-of-Experts model built on the AfmoeForCausalLM architecture, featuring approximately 398 billion total parameters with about 13 billion active parameters per token. Designed as a frontier agentic model, it is specifically engineered for long-horizon planning, multi-step tool calling, and complex reasoning workflows. The model distinguishes itself by generating explicit reasoning traces within dedicated tags before providing a final response, a process that allows it to maintain coherence and accuracy during intricate, multi-turn interactions.
Developed through an intensive two-month scaling effort, the model utilizes an advanced pipeline of supervised fine-tuning and reinforcement learning applied to the Trinity-Large-Base checkpoint. This post-training lineage enables the model to excel in agentic benchmarks, including high performance on LiveCodeBench and τ²-Bench. By offering downloadable weights under an Apache 2.0 license, the model provides a sovereign, auditable alternative for enterprises seeking to avoid vendor lock-in while maintaining the flexibility to deploy on-premises or via cloud APIs for demanding, production-grade agentic applications.