Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Arcee logo

Model details

Trinity Large Thinking

Trinity-Large-Thinking is a reasoning-optimized variant within Arcee's Trinity-Large family, structured as a 398B-parameter sparse Mixture-of-Experts model with roughly 13B active parameters per token. It is built on Trinity-Large-Base and refined through extended chain-of-thought post-training combined with agentic reinforcement learning. This lineage gives it a foundation tuned for deliberate multi-step reasoning rather than just single-pass generation.

The model is designed primarily for agentic workloads, where it can plan, call tools, and sustain multi-turn reasoning through explicit reasoning traces wrapped in dedicated blocks that must remain in context throughout a conversation. Reported benchmark results include 94.7% on τ²-Bench, 91.9% on PinchBench, and 98.2% on LiveCodeBench, positioning it as a strong fit for coding assistants, workflow automation, and complex tool-using agents that benefit from open-weight deployment and transparent intermediate reasoning.

Arceetrinity-large-thinkingtrinitybeta

Quick Info

Powered by
Provider
Arcee
Model key
trinity-large-thinking
Release date
Apr 1, 2026
Last updated
May 28, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.80

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Transparent token rates

Compare Trinity Large Thinking pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Trinity Large Thinking

Arcee

CoverageBenchmark

OpenRouter lists Trinity Large Thinking as an open-weight reasoning model from Arcee AI optimized for agentic workflows, with strong performance on PinchBench and agentic workloads. The model is released on OpenRouter with a 262K token context window and a free tier that is rate-limited, alongside a paid tier priced at As of the listing date, the model shows insufficient traffic data to display token volume or request metrics on OpenRouter. The documentation recommends preserving reasoning tokens in multi-turn conversations and agentic loops to maintain performance, linking to best practices guides for reasoning token handling. The f

Arcee

CoverageBenchmark

Artificial Analysis provides independent evaluation of Trinity Large Thinking, placing it at an Intelligence Index score of 19 out of 111, which is below the median of 29 among comparable large open-weight reasoning models. The model demonstrates notably fast inference at 311 output tokens per second, with API pricing Technical specifications confirm 512K context window (equivalent to 768 pages of size 12 Arial font), 399B total parameters with 13B active per token, text input and output modalities, and Apache 2.0 license with weights available on Hugging Face. The analysis describes Trinity Large Thinking as "below average in intel

Videos about Trinity Large Thinking

More models around Trinity Large Thinking