Currently listed through these providers:
Model details
Inkling
Inkling is an open-weights foundation model released by Thinking Machines Lab, positioned as a flexible base that people can adapt to their own needs rather than as the strongest model on the market. It is a Mixture-of-Experts transformer with 975B total parameters and 41B active per inference, designed to balance capability with efficiency. Training was deliberately broad and balanced: the model was pretrained on 45 trillion tokens spanning text, images, audio, and video, which gives it native multimodal reasoning across those domains and supports a context window of up to one million tokens for long, cross-format workflows.
Inkling represents the first entry in a planned family of models at different sizes, with a lighter Inkling-Small preview shared alongside it for lower-latency and lower-cost use cases. The full model emphasizes controllable thinking effort, multimodal understanding, and open-weight availability as practical strengths for teams that want to fine-tune, deploy privately, or experiment without depending on a closed API. Its fit lies in being a broad, customizable starting point: well suited to projects that need an adaptable open base with strong multimodal grounding and the ability to scale down to a smaller sibling when efficiency matters more than peak capability.
Quick Info
Powered by- Provider
- Nvidia
- Model key
- thinkingmachines/inkling
- Release date
- Jul 15, 2026
- Last updated
- Jul 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
A provider subscription or plan supersedes token-based pricing for this model.
Limits
- Output tokens
- 16,384 tokens
- Context window
- 1,048,576 tokens
Latest news about Inkling
Videos about Inkling
Recent tweets and retweets from Nvidia
More models around Inkling
This exact model name is also listed by 19 other providers.