Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Requesty logo

Model details

Inkling

Inkling is a multimodal autoregressive Mixture-of-Experts transformer with 975 billion total parameters, of which roughly 41 billion are active for any given inference pass. It was trained from scratch by Thinking Machines Lab on a broad 45-trillion-token mixture spanning text, images, audio, and video, and the full weights were released under the Apache 2.0 license so researchers and developers can fine-tune or adapt the model for their own products. The companion announcement also previews Inkling-Small, a lighter variant trained with a similar recipe that targets lower latency and cost while keeping strong general performance, signaling that the team intends to grow a family of differently sized open models rather than ship a single flagship.

In practice the model is aimed at developers building agentic and tool-use systems, coding assistants, retrieval-augmented generation pipelines, chatbots, and other instruction-following applications in English and additional languages. Its native reasoning over text, images, and audio, combined with controllable thinking effort, positions it as a flexible foundation rather than a leaderboard-focused system: the provider is explicit that stronger closed and open models exist, but argues that the combination of multimodal input handling, efficient reasoning, and full weight availability makes Inkling a useful base for downstream customization. A very long context window of up to one million tokens further supports tasks such as large codebase analysis, long-document question answering, and multi-step agent workflows where extended working memory matters.

Requestyinklingling

Quick Info

Powered by
Provider
Requesty
Model key
inkling
Release date
Jul 15, 2026
Last updated
Jul 15, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.87
Output token cost
$4.68

Limits

Output tokens
32,768 tokens
Context window
65,536 tokens

Transparent token rates

Compare Inkling pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inkling

Requesty

CoverageBenchmark

A Medium explainer by TONI RAMCHANDANI frames Thinking Machines Lab's Inkling as a 975-billion-parameter model trained from scratch on 45 trillion multimodal tokens, with creators openly conceding that it is not the strongest model available either open or closed. The piece emphasizes that Inkling's significance lies l The article adds context on Inkling's "controllable reasoning" effort-setting mechanism and discusses practical questions about openness, the memory footprint implied by the total parameter count, and developer cost considerations, while stopping short of first-party technical evidence on the Requesty-routed variant. T

Requesty

CoverageBenchmark

Thinking Machines Lab released Inkling, a 975B-parameter open-weight Mixture-of-Experts model that activates 41B parameters per token and supports a context window of up to 1,048,576 tokens, placing it in the same general size class as Kimi K2.5 and GLM-5.2. The decoder stack has 66 layers with a hidden size of 6,144; Inkling is a native multimodal model that routes images and video frames through a four-layer hMLP and audio through Thinking Machines Lab's dMel representation, with all modalities entering the same text-outputting decoder. Attention is split into 55 sliding-window layers and 11 global layers in a repeating 5-to-1 pat

Videos about Inkling

More models around Inkling