Trinity Large Thinking is a 398-billion parameter sparse Mixture-of-Experts model designed for complex, multi-turn agentic workflows. Rather than activating all parameters for every token, it selectively engages only about 13 billion parameters at inference time—a design choice that preserves the model's raw reasoning power while keeping computational costs manageable. The architecture stacks a hybrid sliding-window and full-attention mechanism across 60 layers, giving the model the ability to handle both localized patterns and long-range dependencies within its 256,000-token context window. Released under the permissive Apache 2.0 license, it is built to be both a capable reasoning engine and an accessible foundation for developers who want to fine-tune or quantize a frontier-class model on their own infrastructure.
The model emerged from Arcee AI, a San Francisco-based lab that took the unusual step of training a foundation model from scratch rather than fine-tuning an existing one. That bet was substantial: nearly half of the roughly $50M in total funding raised was committed to a single 33-day training run on a cluster of 2,048 NVIDIA B300 Blackwell GPUs, a newer generation roughly twice as fast as the prior Hopper architecture. On PinchBench—a comprehensive benchmark measuring real-world agent performance—Trinity Large Thinking ranks second among all existing models, narrowly trailing only Claude Opus-4.6 while delivering results at a fraction of the cost. Its strong showing on SWE-bench Verified and Tau2-Bench Airline further highlights its suitability for coding tasks and structured airline operations. Developers can deploy it through OpenRouter or direct API access, and because the weights are fully open, it invites experimentation in fine-tuning, distillation, and custom agent pipelines.