Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Impossibl logo

Model details

Inkling

Inkling is a multimodal autoregressive transformer from Thinking Machines Lab, Inc., released as an open-weight model under the permissive Apache 2.0 license so developers can download, fine-tune, and self-host it. Independent reporting indicates its mixture-of-experts design follows DeepSeek-V3, while the post-training cold start relied on synthetic data generated by open-weight models such as Kimi K2.5, framing Inkling as an engineering convergence of publicly documented research rather than a from-scratch design. The official model card positions it as a general-purpose system that accepts text, image, and audio inputs and produces text outputs, aimed at developers building agentic tools, coding assistants, chatbots, and retrieval-augmented generation pipelines.

In qualitative benchmark behavior, Inkling trails leading Chinese open models like Kimi and GLM on comprehensive reasoning, programming, and agent evaluations, yet it shows relative strengths in web application design, audio understanding, safety testing, and certain math tasks. Rather than positioning the model as a benchmark leader, Thinking Machines is emphasizing enterprise customization through its Tinker platform, letting organizations adapt Inkling into domain-specific assistants with the compliance benefits of a US-developed open-weight deployment option.

Impossiblthinkingmachines/inklingling

Quick Info

Powered by
Provider
Impossibl
Model key
thinkingmachines/inkling
Release date
Jul 15, 2026
Last updated
Jul 15, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.87
Output token cost
$4.68

Limits

Output tokens
65,536 tokens
Context window
65,536 tokens

Transparent token rates

Compare Inkling pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inkling

Vercel AI Gateway

Official sourceAnnouncement

Thinking Machines Lab released Inkling on July 15, 2026, as its first in-house model and the inaugural entry in a planned model family. The official announcement describes Inkling as a Mixture-of-Experts transformer with 975 billion total parameters and 41 billion active per token, pretrained on 45 trillion tokens span Alongside the main release, the announcement previews Inkling-Small, a lighter-weight variant with 12 billion active parameters trained with a similar recipe, aimed at lower cost and latency. The post also introduces an Inkling Playground in the Tinker console for chatting with the model, and demonstrates the model wri

Thinking Machines

Coverage

Thinking Machines Lab released Inkling on July 15, 2026, and the model adopts a Mixture-of-Experts architecture with 975 billion total parameters and 41 billion activated per token, according to a 36Kr report. Pre-training used 45 trillion tokens spanning text, images, audio, and video, and the model supports a maximum The same report notes that Inkling's weights were published under an Apache 2.0 license, enabling self-deployment, and that developers can fine-tune the model through Thinking Machines' Tinker platform. The article also describes the architecture as following the DeepSeek-V3 MoE design, with post-training cold-start da

Vercel AI Gateway

Coverage

A detailed third-party breakdown corroborates and expands on the official Inkling launch with operational context for teams evaluating self-hosting. It confirms the core spec sheet — 975B total / 41B active MoE, 45T multimodal training tokens, Apache 2.0 license, full weights on Hugging Face — and notes a practical dis The breakdown also covers the controllable thinking-effort dial and the practical infrastructure considerations of self-hosting a model at this scale, useful for developers and platform teams weighing procurement and compliance. It highlights Tinker as the same-day fine-tuning channel, the self-finetuning demonstration

Impossibl

Coverage

A Hugging Face community blog post (co-published with Thinking Machines) details day-0 developer tooling for Inkling: BF16 and NVFP4 weight variants, speculative multi-token prediction (MTP) layers for faster inference, and first-day integration in transformers, SGLang, vLLM, and llama.cpp. The post frames Inkling as a The same post was later updated to cover the release of Inkling-Small (276B total / 12B active parameters) and the Inkling-Small-NVFP4 variant, plus MXFP8 weight support, one-click deployment on Hugging Face Inference Endpoints (up to 160 TPS for Inkling-Small), and a real-time voice and image demo. A linked collection

Videos about Inkling

More models around Inkling