Currently listed through these providers:
Model details
Inkling
Inkling is an open-weights foundation model released by Thinking Machines Lab as a base for further customization rather than a claim of overall leadership among available models. It is built as a Mixture-of-Experts transformer with 975B total parameters and 41B active parameters per inference, paired with a context window of up to one million tokens. The model was pretrained from scratch on roughly 45 trillion tokens spanning text, images, audio, and video, giving it native multimodal reasoning over those inputs while emitting text. The design intent, as described by the lab, is to combine broad coverage across domains with efficient and controllable thinking effort, so users can tune the balance between cost and quality for their own applications.
In practical terms, Inkling suits teams that want a large open-weights model they can fine-tune or adapt rather than consume only through a fixed API. Its sparse activation keeps per-token compute manageable despite the very large total capacity, while the long context window makes it relevant for document-heavy, code-base, or multimodal transcript workflows. Thinking Machines Lab has positioned Inkling as the first in a family of models, with a lighter-weight preview called Inkling-Small already shared alongside it, signaling a roadmap of size variants trained with a similar recipe so users can pick the scale that matches their latency and cost constraints.
Quick Info
Powered by- Provider
- OpenRouter
- Model key
- thinkingmachines/inkling
- Release date
- Jul 15, 2026
- Last updated
- Jul 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.00
- Output token cost
- $4.05
Limits
- Output tokens
- 471,859 tokens
- Context window
- 1,048,576 tokens
Latest news about Inkling
Videos about Inkling
More models around Inkling
This exact model name is also listed by 19 other providers.