Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Baseten logo

Model details

Inkling Small

Inkling Small is the compact member of Thinking Machines Lab's Inkling family, built as an efficient Mixture-of-Experts transformer with 276B total parameters and 12B active per token. It uses a 42-layer decoder with a hidden size of 4096, hybrid attention made up of a 512-token local sliding window with every sixth layer attending globally, and the same sparse feed-forward routing as the larger Inkling checkpoint. The model keeps the family's multimodal backbone, so it accepts text, images, and audio, and produces text outputs, giving it the visual and acoustic reasoning capabilities that distinguish the Inkling line from text-only peers.

Thinking Machines Lab positions Inkling Small as a way to get Inkling-class reasoning at roughly a quarter of the size, and the release highlights competitive or superior results against peers in its weight class on agentic tool use (Terminal-Bench 2.1), reasoning (HLE text-only), and instruction following (IFBench). It supports variable thinking effort so users can trade compute for quality on a per-task basis, and it runs within a one-the cataloged API limit that suits long documents, extended conversations, and tool-driven workflows. Trained on GB300 NVL72 systems and shipped as an open-weights release, it is a practical choice for teams that want strong multimodal reasoning without paying for the full 975B Inkling footprint.

Basetenthinkingmachines/inkling-smallling

Quick Info

Powered by
Provider
Baseten
Model key
thinkingmachines/inkling-small
Release date
Jul 30, 2026
Last updated
Jul 30, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.50
Output token cost
$1.20

Limits

Output tokens
32,768 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Inkling Small pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inkling Small

OpenRouter

Official sourceAnnouncement

Thinking Machines Lab's official launch post introduces the Inkling family and explicitly previews Inkling-Small as a lighter-weight sibling with 12 billion active parameters, trained with a similar recipe to the flagship Inkling. The post confirms the model is available immediately for fine-tuning on the company's Tin The announcement positions Inkling-Small as a customization-friendly base model that achieves strong performance at lower cost and latency than the 975B-parameter flagship, which uses 41B active parameters with a 1M token context. Both models share native multimodal reasoning across text, images and audio with controll

Baseten

Coverage

Silicon Republic's July 31, 2026 coverage reports that Thinking Machines Lab released Inkling-Small at 276 billion total parameters with 12 billion active, achieving what the company describes as comparable or better performance than the original 975B-parameter Inkling on reasoning and agentic tasks. Artificial Analysi Thinking Machines reports Inkling-Small scored above 31% on Humanity's Last Exam (versus Inkling's 29.7%) and crossed 80% on SWE-Bench Verified. The model accepts text, image, and audio inputs and produces text outputs, with Thinking Machines highlighting its use for cropping, zooming, and programmatic image inspection

Pioneer

CoverageBenchmark

VentureBeat reports that Thinking Machines released Inkling-Small, an open-source Apache 2.0–licensed multimodal reasoning model positioned as a smaller sibling to its flagship Inkling. Per the article, Inkling-Small has 276 billion total parameters with 12 billion active parameters per token, accepts text, image, and According to VentureBeat, Artificial Analysis assigned Inkling-Small a score of 40 on its Intelligence Index, within a single point of the larger flagship Inkling's score of 41 despite the smaller model being roughly one-quarter the size. The article frames Inkling-Small as a fit for enterprises with moderate GPU resou

OpenRouter

Coverage

Thinking Machines Lab released an updated Hugging Face blog post confirming that Inkling-Small is now available as a companion to the full Inkling model. The post explicitly names Inkling-Small with 276 billion total parameters and 12 billion active parameters, sharing the same architecture as the larger Inkling, and a Day-0 runtime coverage is confirmed across transformers, SGLang, vLLM and llama.cpp for Inkling-Small, and a real-time voice and image demo is shipped so developers can interact with the variant directly. The post positions Inkling-Small as part of a broader collection of Inkling models and emphasizes its role as a lig

Baseten

CoverageBenchmark

OpenRouter's listing identifies Thinking Machines' Inkling Small as an open-weight multimodal mixture-of-experts model with 12B active parameters out of 276B total, supporting a 1M-token context window, multimodal inputs (text, image, audio) with text output, and a release date of July 30, 2026 — the same day Baseten's The free Inkling Small endpoint is gated to agentic harnesses under the TML Free Research API Terms of Service, with session prompts and outputs logged and disassociated from account identifiers for model improvement — a caveat that does not necessarily apply to Baseten's commercial Model API endpoint. Latency on the f

Baseten

CoverageBenchmark

Artificial Analysis lists Inkling Small as one of 20 models offered on the platform, ranking it as the fastest model on Baseten at 336 output tokens per second, ahead of GLM-5.2 (max) (FAST) at 283 t/s, gpt-oss-120b (low) at 276 t/s, and the larger Inkling at 247 t/s. On price, Inkling Small sits at a blended $0.29 per The page also reports that Inkling Small supports a 1M-token context window alongside Kimi K3 (max), GLM-5.2 (max), and DeepSeek V4 Flash 0731 (max), the largest context tiers on Baseten. Artificial Analysis Intelligence Index v4.1.1 — composed of GDPval-AA v2, τ³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last

Baseten

Official sourceRelease Notes

Baseten's official changelog confirms that Inkling Small became available through Baseten Model APIs on July 30, 2026, served as thinkingmachines/inkling-small through an OpenAI-compatible endpoint using a Baseten API key, with dedicated deployments also referenced. The same changelog documents Basete's continued platf For developers, the changelog entry signals that Inkling Small can be invoked through the standard Baseten OpenAI-compatible endpoint with the thinkingmachines/inkling-small model identifier, without bespoke integration work, and that dedicated capacity is available for production use cases. Adjacent changelog items —

Videos about Inkling Small

More models around Inkling Small