Sulat.com
AI models
Charm Hyper logo

Model details

Inkling

Inkling is a one-trillion-parameter multimodal model trained by Thinking Machines Lab, designed to take in text, image, and audio inputs and produce text output. The model introduces a distinctive architectural blend that combines relative attention, short convolution, and shared expert sinks, giving it an unusually wide context reach of up to one million tokens. Open-weight checkpoints for Inkling are published on Hugging Face in both BF16 and NVFP4 formats, making it straightforward for teams to self-host and fine-tune on their own infrastructure. The combination of a very long effective context window and native multimodal intake positions Inkling as a strong fit for tasks that require reasoning over long documents, mixed media, or extended conversational histories.

Day-zero integration with vLLM brings practical performance and serving flexibility to Inkling deployments. On four GB200 GPUs, the model can reach up to 380 tokens per second per user with multi-token prediction enabled, or around 140 tokens per second without it. vLLM provides full feature parity for the run, including LoRA adapters, tensor, data, expert, and pipeline parallelism, prefix caching, and disaggregated serving, which lets operators scale Inkling across clusters while keeping throughput efficient. Google Cloud also lists Inkling support within Vertex AI Model Garden, giving teams an additional managed path to deploy the model. Together, these integrations make Inkling approachable for production workloads that need both high-speed inference and the ability to handle very long, multimodal inputs.

Charm Hyperinklingling

Quick Info

Powered by
Provider
Charm Hyper
Model key
inkling
Release date
Aug 15, 2026
Last updated
Sep 5, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$1.0888
Output token cost
$4.40964

Limits

Output tokens
32,768 tokens
Context window
1,048,576 tokens

Transparent token rates

Compare Inkling pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inkling

No articles yet. Fetch the latest news to show it here.

Videos about Inkling

More models around Inkling