Currently listed through these providers:
Model details
Inkling
Inkling is designed as a general-purpose multimodal autoregressive transformer for applications that combine text with images or audio and produce text responses. Its documented architecture is a 66-layer decoder-only model, and the available model-card and catalog descriptions support its use across English, other languages, and multiple programming languages. Distributed builds include BF16 and NVFP4 variants, with the NVFP4 artifacts shown as signed in the catalog, giving developers options for different deployment requirements.
The model is intended for developers building coding assistants, agentic and tool-use systems, chatbots, retrieval-augmented generation systems, instruction-following applications, and general conversational or multimodal experiences. That combination makes it a practical fit when one model must handle varied inputs and text-based interaction rather than being limited to a single task. The available evidence does not provide supported benchmark results or verified training-scale figures, so comparisons should focus on its documented use cases and deployment format.
Quick Info
Powered by- Provider
- Modal
- Model key
- thinkingmachines/Inkling-NVFP4
- Release date
- Jul 15, 2026
- Last updated
- Jul 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.20
- Output token cost
- $5.00
Limits
- Output tokens
- 262,144 tokens
- Context window
- 1,048,576 tokens
Transparent token rates
Compare Inkling pricing
Rates are shown per one million tokens. Combined means one million input plus one million output tokens.
Latest news about Inkling
No articles yet. Fetch the latest news to show it here.
Videos about Inkling
More models around Inkling
This exact model name is also listed by 21 other providers.