Currently listed through these providers:
Model details
Inkling (Fireworks AI)
Inkling is an open-weights foundation model from Thinking Machines Lab, built as a broad generalist intended for fine-tuning across a wide range of downstream tasks. Its architecture is a 975-billion-parameter sparse Mixture-of-Experts design with 41 billion active parameters per forward pass, combining hybrid local and global attention to balance long-context efficiency with fine-grained reasoning. The model was trained natively across text, image, and audio, positioning it as a multimodal audio-and-vision foundation rather than a text-only system, and it exposes controllable thinking effort so developers can tune how deeply the model reasons per query.
Practically, Inkling is suited for teams that need very long-context reasoning and multimodal inputs in a single checkpoint, with a one-million-token context window that makes it a fit for document-heavy workflows, codebases, or extended conversational sessions. Its combination of sparse activation, hybrid attention, and adjustable reasoning depth suggests a design aimed at flexible deployment rather than narrow specialization, allowing practitioners to dial compute up for complex analysis or down for lightweight inference. For organizations exploring open-weight alternatives for fine-tuning at scale, Inkling represents one of the larger publicly released MoE checkpoints available.
Quick Info
Powered by- Provider
- Eden AI
- Model key
- fireworks_ai/accounts/fireworks/models/inkling
- Release date
- Jul 15, 2026
- Last updated
- Jul 15, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $1.00
- Output token cost
- $4.05
Limits
- Output tokens
- 1,048,576 tokens
- Context window
- 1,048,576 tokens
Latest news about Inkling (Fireworks AI)
No articles yet. Fetch the latest news to show it here.
Videos about Inkling (Fireworks AI)
More models around Inkling (Fireworks AI)
This exact model name is also listed by 21 other providers.