DevPass (LLM Gateway)
Sebastian Raschka's architecture analysis details that Inkling is a 975B-parameter open-weight Mixture-of-Experts model that activates 41B parameters per token and supports a context window of up to 1,048,576 tokens, placing it in the same general size class as Kimi K2.5 and GLM-5.2. The model has 66 decoder layers wit The architecture introduces several less-common details: 55 sliding-window attention layers and 11 global attention layers in a repeating 5-to-1 local-global pattern, with local layers using a 512-token window and 4-to-1 grouped-query attention while global layers use 8-to-1 GQA. Inkling skips RoPE in favor of a learne