Currently listed through these providers:
Model details
Inkling
Inkling is a one-trillion-parameter multimodal model trained by Thinking Machines Lab, designed to take in text, image, and audio inputs and produce text output. The model introduces a distinctive architectural blend that combines relative attention, short convolution, and shared expert sinks, giving it an unusually wide context reach of up to one million tokens. Open-weight checkpoints for Inkling are published on Hugging Face in both BF16 and NVFP4 formats, making it straightforward for teams to self-host and fine-tune on their own infrastructure. The combination of a very long effective context window and native multimodal intake positions Inkling as a strong fit for tasks that require reasoning over long documents, mixed media, or extended conversational histories.
Day-zero integration with vLLM brings practical performance and serving flexibility to Inkling deployments. On four GB200 GPUs, the model can reach up to 380 tokens per second per user with multi-token prediction enabled, or around 140 tokens per second without it. vLLM provides full feature parity for the run, including LoRA adapters, tensor, data, expert, and pipeline parallelism, prefix caching, and disaggregated serving, which lets operators scale Inkling across clusters while keeping throughput efficient. Google Cloud also lists Inkling support within Vertex AI Model Garden, giving teams an additional managed path to deploy the model. Together, these integrations make Inkling approachable for production workloads that need both high-speed inference and the ability to handle very long, multimodal inputs.
Quick Info
Powered by- Provider
- Charm Hyper
- Model key
- inkling
- Release date
- Aug 15, 2026
- Last updated
- Sep 5, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.0888
- Output token cost
- $4.40964
Limits
- Output tokens
- 32,768 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Inkling pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Inkling
No articles yet. Fetch the latest news to show it here.
Videos about Inkling
More models around Inkling
This exact model name is also listed by 20 other providers.