Inkling Small is the compact member of Thinking Machines Lab's Inkling family, built as an efficient Mixture-of-Experts transformer with 276B total parameters and 12B active per token. It uses a 42-layer decoder with a hidden size of 4096, hybrid attention made up of a 512-token local sliding window with every sixth layer attending globally, and the same sparse feed-forward routing as the larger Inkling checkpoint. The model keeps the family's multimodal backbone, so it accepts text, images, and audio, and produces text outputs, giving it the visual and acoustic reasoning capabilities that distinguish the Inkling line from text-only peers.
Thinking Machines Lab positions Inkling Small as a way to get Inkling-class reasoning at roughly a quarter of the size, and the release highlights competitive or superior results against peers in its weight class on agentic tool use (Terminal-Bench 2.1), reasoning (HLE text-only), and instruction following (IFBench). It supports variable thinking effort so users can trade compute for quality on a per-task basis, and it runs within a one-the cataloged API limit that suits long documents, extended conversations, and tool-driven workflows. Trained on GB300 NVL72 systems and shipped as an open-weights release, it is a practical choice for teams that want strong multimodal reasoning without paying for the full 975B Inkling footprint.