Currently listed through these providers:
Model details
Inkling Small
Inkling Small is part of Thinking Machines Lab's Inkling family, previewed alongside the flagship Inkling release as a lighter-weight variant designed for lower cost and latency while retaining a similar training approach. The flagship Inkling is described as a Mixture-of-Experts transformer with 975 billion total parameters and 41 billion active, pretrained on 45 trillion tokens of text, images, audio, and video, and Inkling Small shares that lineage with 12 billion active parameters under a comparable recipe. The model is published with open weights on Hugging Face under the thinkingmachines organization in BF16 and NVFP4 formats, making it a flexible foundation for downstream customization and fine-tuning.
Designed as a general-purpose multimodal system, Inkling Small accepts text, image, and audio inputs and produces text outputs, supporting English plus broader multilingual use. The model card positions it for developers building agentic and tool-use systems, coding assistants, chatbots, retrieval-augmented generation pipelines, and other conversational or instruction-following applications. Its intended audience, combined with native reasoning over multiple modalities and open-weight availability, makes it a practical base for teams that want to adapt a capable multimodal model rather than rely on a closed endpoint.
Quick Info
Powered by- Provider
- LLMTR
- Model key
- thinkingmachines/inkling-small
- Release date
- Jul 30, 2026
- Last updated
- Jul 30, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.58
- Output token cost
- $1.44
Limits
- Output tokens
- 262,144 tokens
- Context window
- 262,144 tokens
Latest news about Inkling Small
Videos about Inkling Small
More models around Inkling Small
This exact model name is also listed by 9 other providers.