Inkling Small free is the no-cost variant of Thinking Machines' smaller Inkling model, an open-weight multimodal mixture-of-experts design that activates 12B parameters out of a much larger 276B total. That MoE architecture lets it handle reasoning, coding, agentic workflows, retrieval-augmented generation, instruction following, and multilingual conversation while keeping per-query compute relatively modest, making it a sensible pick when you want Inkling-quality behavior without paying per-token fees. Because it is open-weight, teams can also self-host or audit behavior rather than treating the system as a black box, which matters for developer workflows that need reproducibility or fine-grained control over prompts and context.
On independent benchmarks, Inkling Small free matches its paid sibling in core intelligence measures, posting a 41.2 Intelligence Index, a 52.9 Coding Index, and 90% on GPQA Diamond, so the zero-cost tier does not force a quality sacrifice on the tasks most teams care about. The trade-off shows up in throughput and responsiveness: the free variant runs at roughly 104 tokens per second with around 913 ms latency, noticeably slower than the paid edition's 175 tokens per second and 237 ms, which is worth weighing for latency-sensitive agentic loops or interactive coding sessions. With a very large context window, multimodal inputs, and tool-calling plus reasoning support exposed through Kilo's zero-markup gateway, it is well suited as a default free workhorse model for experimentation, prototyping, and steady development workloads where some extra waiting time is acceptable.