Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Vercel AI Gateway logo

Model details

Inkling Small

Inkling Small is the compact member of Thinking Machines Lab's Inkling family, built as an efficient Mixture-of-Experts transformer with 276B total parameters and 12B active per token. It uses a 42-layer decoder with a hidden size of 4096, hybrid attention made up of a 512-token local sliding window with every sixth layer attending globally, and the same sparse feed-forward routing as the larger Inkling checkpoint. The model keeps the family's multimodal backbone, so it accepts text, images, and audio, and produces text outputs, giving it the visual and acoustic reasoning capabilities that distinguish the Inkling line from text-only peers.

Thinking Machines Lab positions Inkling Small as a way to get Inkling-class reasoning at roughly a quarter of the size, and the release highlights competitive or superior results against peers in its weight class on agentic tool use (Terminal-Bench 2.1), reasoning (HLE text-only), and instruction following (IFBench). It supports variable thinking effort so users can trade compute for quality on a per-task basis, and it runs within a one-the cataloged API limit that suits long documents, extended conversations, and tool-driven workflows. Trained on GB300 NVL72 systems and shipped as an open-weights release, it is a practical choice for teams that want strong multimodal reasoning without paying for the full 975B Inkling footprint.

Vercel AI Gatewaythinkingmachines/inkling-smallling

Quick Info

Powered by
Provider
Vercel AI Gateway
Model key
thinkingmachines/inkling-small
Release date
Jul 30, 2026
Last updated
Jul 30, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.45
Output token cost
$1.20

Limits

Output tokens
1,000,000 tokens
Context window
1,000,000 tokens

Transparent token rates

Compare Inkling Small pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inkling Small

OpenRouter

Official sourceAnnouncement

Thinking Machines Lab's official launch post introduces the Inkling family and explicitly previews Inkling-Small as a lighter-weight sibling with 12 billion active parameters, trained with a similar recipe to the flagship Inkling. The post confirms the model is available immediately for fine-tuning on the company's Tin The announcement positions Inkling-Small as a customization-friendly base model that achieves strong performance at lower cost and latency than the 975B-parameter flagship, which uses 41B active parameters with a 1M token context. Both models share native multimodal reasoning across text, images and audio with controll

Pioneer

CoverageBenchmark

VentureBeat reports that Thinking Machines released Inkling-Small, an open-source Apache 2.0–licensed multimodal reasoning model positioned as a smaller sibling to its flagship Inkling. Per the article, Inkling-Small has 276 billion total parameters with 12 billion active parameters per token, accepts text, image, and According to VentureBeat, Artificial Analysis assigned Inkling-Small a score of 40 on its Intelligence Index, within a single point of the larger flagship Inkling's score of 41 despite the smaller model being roughly one-quarter the size. The article frames Inkling-Small as a fit for enterprises with moderate GPU resou

OpenRouter

Coverage

Thinking Machines Lab released an updated Hugging Face blog post confirming that Inkling-Small is now available as a companion to the full Inkling model. The post explicitly names Inkling-Small with 276 billion total parameters and 12 billion active parameters, sharing the same architecture as the larger Inkling, and a Day-0 runtime coverage is confirmed across transformers, SGLang, vLLM and llama.cpp for Inkling-Small, and a real-time voice and image demo is shipped so developers can interact with the variant directly. The post positions Inkling-Small as part of a broader collection of Inkling models and emphasizes its role as a lig

Videos about Inkling Small

More models around Inkling Small