Sulat.com
AI models
LLMTR logo

Model details

Inkling

Thinking Machines Lab built Inkling from scratch as an open-weights foundation model designed for customization rather than as a leaderboard champion. The team trained it on 45 trillion tokens spanning text, images, audio, and video, and shaped it to reason natively across text, images, and audio while supporting controllable thinking effort so users can balance quality against latency and cost. Because the full weights are released, the model is positioned as a flexible base that organizations, researchers, and developers can adapt through fine-tuning or downstream pipelines rather than consume as a fixed service.

Under the hood, Inkling is a Mixture-of-Experts transformer with 975 billion total parameters and 41 billion active per token, paired with a context window that scales up to one million tokens for long-form reasoning, document analysis, and multi-turn workflows. It arrives as the first entry in a planned family, accompanied by a lighter Inkling-Small preview at 12 billion active parameters that shares the same architectural philosophy, signaling a roadmap toward scaled variants optimized for different deployment envelopes. Practical strengths therefore center on broad multimodal grounding, efficient expert routing, large-context retention, and the freedom that open weights provide for private fine-tuning and domain-specific adaptation.

LLMTRthinkingmachines/inklingling

Quick Info

Powered by
Provider
LLMTR
Model key
thinkingmachines/inkling
Release date
Jul 15, 2026
Last updated
Jul 15, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.87
Output token cost
$4.68

Limits

Output tokens
262,144 tokens
Context window
262,144 tokens

Latest news about Inkling

Videos about Inkling

More models around Inkling