Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Trinity Large Thinking

Trinity Large Thinking is a sparse Mixture-of-Experts reasoning model built around an AfmoeForCausalLM architecture, totaling roughly 400 billion parameters with about 13 billion active per token during inference. The model is explicitly designed for long-horizon planning, tool use, and multi-step agent workflows, emitting structured reasoning traces inside `<think ...` blocks so developers can inspect how conclusions were reached. Its design as a frontier-class open-weight model reflects Arcee's intent to offer a Western alternative to the Chinese open models that have dominated the reasoning space, enabling both on-premises deployments and API access for teams who want full control over their inference environment.

The model was post-trained using extended chain-of-thought reasoning and agentic reinforcement learning to cultivate the kind of multi-turn deliberation that real-world agent tasks demand. Multiple GPU configurations are documented in the wild—vLLM recipes reference tensor-parallel setups across 8 GPUs for production latency, and community discussions suggest it can compress into two DGX Spark nodes. Quantized GGUF variants broaden accessibility for users running on constrained hardware, while the Apache 2.0 license removes deployment barriers for enterprises building autonomous workflows. The result is a model tuned for agents that must plan several steps ahead, call external tools reliably, and maintain coherent context across extended conversations.

Kilo Gatewayarcee-ai/trinity-large-thinkingtrinity

Quick Info

Powered by
Provider
Kilo Gateway
Model key
arcee-ai/trinity-large-thinking
Release date
Apr 1, 2026
Last updated
May 28, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.25
Output token cost
$0.80

Limits

Output tokens
80,000 tokens
Context window
262,144 tokens

Transparent token rates

Compare Trinity Large Thinking pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Trinity Large Thinking

No articles yet. Fetch the latest news to show it here.

Videos about Trinity Large Thinking

More models around Trinity Large Thinking