Currently listed through these providers:
Model details
Inkling
Thinking Machines Lab built Inkling from scratch as an open-weights foundation model designed for customization rather than as a leaderboard champion. The team trained it on 45 trillion tokens spanning text, images, audio, and video, and shaped it to reason natively across text, images, and audio while supporting controllable thinking effort so users can balance quality against latency and cost. Because the full weights are released, the model is positioned as a flexible base that organizations, researchers, and developers can adapt through fine-tuning or downstream pipelines rather than consume as a fixed service.
Under the hood, Inkling is a Mixture-of-Experts transformer with 975 billion total parameters and 41 billion active per token, paired with a context window that scales up to one million tokens for long-form reasoning, document analysis, and multi-turn workflows. It arrives as the first entry in a planned family, accompanied by a lighter Inkling-Small preview at 12 billion active parameters that shares the same architectural philosophy, signaling a roadmap toward scaled variants optimized for different deployment envelopes. Practical strengths therefore center on broad multimodal grounding, efficient expert routing, large-context retention, and the freedom that open weights provide for private fine-tuning and domain-specific adaptation.
Quick Info
Powered by- Provider
- LLMTR
- Model key
- thinkingmachines/inkling
- Release date
- Jul 15, 2026
- Last updated
- Jul 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.87
- Output token cost
- $4.68
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Inkling
Videos about Inkling
More models around Inkling
This exact model name is also listed by 19 other providers.