Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
DevPass (LLM Gateway) logo

Model details

Inkling

The model overview is temporarily unavailable.

DevPass (LLM Gateway)inklingling

Quick Info

Powered by
Provider
DevPass (LLM Gateway)
Model key
inkling
Release date
Jul 15, 2026
Last updated
Jul 15, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.95
Output token cost
$4.05

Limits

Output tokens
1,048,576 tokens
Context window
524,288 tokens

Transparent token rates

Compare Inkling pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about Inkling

DevPass (LLM Gateway)

CoverageBenchmark

Sebastian Raschka's architecture analysis details that Inkling is a 975B-parameter open-weight Mixture-of-Experts model that activates 41B parameters per token and supports a context window of up to 1,048,576 tokens, placing it in the same general size class as Kimi K2.5 and GLM-5.2. The model has 66 decoder layers wit The architecture introduces several less-common details: 55 sliding-window attention layers and 11 global attention layers in a repeating 5-to-1 local-global pattern, with local layers using a 512-token window and 4-to-1 grouped-query attention while global layers use 8-to-1 GQA. Inkling skips RoPE in favor of a learne

DevPass (LLM Gateway)

Coverage

Thinking Machines Lab has released Inkling, described as its first open-weights model and a mixture-of-experts (MoE) multimodal system that natively accepts image and audio inputs alongside text. The model is published under the permissive Apache 2.0 license, one of the least restrictive terms an AI lab can choose, mea The MoE architecture activates only a subset of parameters for any given input, a sparse routing approach that lets teams scale total capacity while keeping per-forward-pass compute cost manageable. Pairing that architecture with native image and audio handling places Inkling in the growing category of open models buil

DevPass (LLM Gateway)

CoverageRelease Notes

Thinking Machines Lab has released Inkling, its first model, as an open-weight system trained from scratch to handle audio, video, and text inputs. WIRED reports the model has 975 billion parameters and is designed for advanced reasoning and coding, though it requires a cluster of specialized chips to run. The lab nota The release positions Thinking Machines Lab — founded by former OpenAI executives including Mira Murati, John Schulman, and Lilian Weng in February 2025 — within the open-source AI race, where it claims Inkling performs at a level comparable to leading open-weight models from China. WIRED frames the open-weight approac

Videos about Inkling

More models around Inkling