North Mini Code marks Cohere's entry into agentic coding and the opening of its North family, pairing a sparse mixture-of-experts design with 30B total parameters and 3B active parameters to keep inference lean. That small active footprint is intended to translate into low-latency responses, including on local hardware, while still handling code generation, multi-step software engineering workflows, and terminal-style tasks. Training emphasizes generalization across agent harnesses such as OpenCode and SWE-Agent, so the same checkpoint can plug into different tool-driven environments without heavy retraining. The model's strengths line up with the workflows developers actually run: the cataloged API limit context window leaves room for large repositories and long debugging sessions, interleaved reasoning is woven into generation so the model can pause to think mid-task, and tool use is exposed through JSON-schema interfaces that fit cleanly with structured-output pipelines. Open weights under Apache 2.0 make it attractive for teams that want to self-host, fine-tune, or audit the model, while zero-cost access through the OpenCode Zen surface removes friction for evaluation. In practice it is best suited to agentic coding setups that need long context, reliable tool calling, and a permissive license, rather than general chat or non-code reasoning workloads.
North Mini Code marks Cohere's entry into agentic coding and the opening of its North family, pairing a sparse mixture-of-experts design with 30B total parameters and 3B active parameters to keep inference lean. That small active footprint is intended to translate into low-latency responses, including on local hardware, while still handling code generation, multi-step software engineering workflows, and terminal-style tasks. Training emphasizes generalization across agent harnesses such as OpenCode and SWE-Agent, so the same checkpoint can plug into different tool-driven environments without heavy retraining. The model's strengths line up with the workflows developers actually run: the cataloged API limit context window leaves room for large repositories and long debugging sessions, interleaved reasoning is woven into generation so the model can pause to think mid-task, and tool use is exposed through JSON-schema interfaces that fit cleanly with structured-output pipelines. Open weights under Apache 2.0 make it attractive for teams that want to self-host, fine-tune, or audit the model, while zero-cost access through the OpenCode Zen surface removes friction for evaluation. In practice it is best suited to agentic coding setups that need long context, reliable tool calling, and a permissive license, rather than general chat or non-code reasoning workloads.