Inkling is a multimodal autoregressive Mixture-of-Experts transformer with 975 billion total parameters, of which roughly 41 billion are active for any given inference pass. It was trained from scratch by Thinking Machines Lab on a broad 45-trillion-token mixture spanning text, images, audio, and video, and the full weights were released under the Apache 2.0 license so researchers and developers can fine-tune or adapt the model for their own products. The companion announcement also previews Inkling-Small, a lighter variant trained with a similar recipe that targets lower latency and cost while keeping strong general performance, signaling that the team intends to grow a family of differently sized open models rather than ship a single flagship.
In practice the model is aimed at developers building agentic and tool-use systems, coding assistants, retrieval-augmented generation pipelines, chatbots, and other instruction-following applications in English and additional languages. Its native reasoning over text, images, and audio, combined with controllable thinking effort, positions it as a flexible foundation rather than a leaderboard-focused system: the provider is explicit that stronger closed and open models exist, but argues that the combination of multimodal input handling, efficient reasoning, and full weight availability makes Inkling a useful base for downstream customization. A very long context window of up to one million tokens further supports tasks such as large codebase analysis, long-document question answering, and multi-step agent workflows where extended working memory matters.