Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Kilo Gateway logo

Model details

Thinking Machines: Inkling Small (free)

Inkling Small free is the no-cost variant of Thinking Machines' smaller Inkling model, an open-weight multimodal mixture-of-experts design that activates 12B parameters out of a much larger 276B total. That MoE architecture lets it handle reasoning, coding, agentic workflows, retrieval-augmented generation, instruction following, and multilingual conversation while keeping per-query compute relatively modest, making it a sensible pick when you want Inkling-quality behavior without paying per-token fees. Because it is open-weight, teams can also self-host or audit behavior rather than treating the system as a black box, which matters for developer workflows that need reproducibility or fine-grained control over prompts and context.

On independent benchmarks, Inkling Small free matches its paid sibling in core intelligence measures, posting a 41.2 Intelligence Index, a 52.9 Coding Index, and 90% on GPQA Diamond, so the zero-cost tier does not force a quality sacrifice on the tasks most teams care about. The trade-off shows up in throughput and responsiveness: the free variant runs at roughly 104 tokens per second with around 913 ms latency, noticeably slower than the paid edition's 175 tokens per second and 237 ms, which is worth weighing for latency-sensitive agentic loops or interactive coding sessions. With a very large context window, multimodal inputs, and tool-calling plus reasoning support exposed through Kilo's zero-markup gateway, it is well suited as a default free workhorse model for experimentation, prototyping, and steady development workloads where some extra waiting time is acceptable.

Kilo Gatewaythinkingmachines/inkling-small:freeling

Quick Info

Powered by
Provider
Kilo Gateway
Model key
thinkingmachines/inkling-small:free
Release date
Jul 30, 2026
Last updated
Jul 30, 2026
Input modalities
Output modalities
Capabilities

Cost

A provider subscription or plan supersedes token-based pricing for this model.

Limits

Output tokens
262,144 tokens
Context window
1,048,576 tokens

Latest news about Thinking Machines: Inkling Small (free)

Kilo Gateway

CoverageBenchmark

Thinking Machines: Inkling Small (free) is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, configured with 12B active parameters out of 276B total, a 1M token context window, multimodal input/output, and a release date of July 30, 2026, per its OpenRouter listing. The model is positioned On the OpenRouter listing, the Thinking Machines Free provider routes the endpoint with a P50 latency of about 1.09 seconds, throughput of roughly 105 tokens per second, and 99.89% uptime over a one-week window; reported weekly throughput percentiles range from a P50 of 99 tok/s up to a P99 of 361 tok/s, and end-to-end

Videos about Thinking Machines: Inkling Small (free)

More models around Thinking Machines: Inkling Small (free)