Nebius Token Factory
We released Qwen3-Next-80B-A3B-Thinking-Uncensored, a large-scale reasoning model where political refusal behavior has been removed while preserving saf...
Model details
Built on the Qwen3-Next architecture, this thinking-focused variant inherits a design that combines hybrid attention with a highly sparse Mixture-of-Experts structure and multi-token prediction. The underlying base checkpoint carries 80 billion total parameters but activates only around 3 billion at inference, a balance that delivers performance comparable to, or slightly better than, the dense Qwen3-32B while using a fraction of the training compute. Reinforcement-learning post-training was applied specifically to stabilize and accelerate reasoning training under the new architecture, producing a model intended for deliberate, multi-step problem solving rather than fast casual chat.
In practice, the model is well suited to long-context analytical workloads such as extended document review, multi-turn research sessions, and agentic pipelines that call external tools, where its sparse activation profile keeps throughput high even past 32K-token contexts. The open weights allow teams to self-host or fine-tune the reasoning checkpoint, and community variants like an "uncensored" build signal an active fine-tuning ecosystem around the release. Independent gateway listings describe it as a chat model with function-calling and structured-output support, making it a flexible drop-in for orchestration layers that need both reasoning depth and reliable tool integration.
Nebius Token Factory
We released Qwen3-Next-80B-A3B-Thinking-Uncensored, a large-scale reasoning model where political refusal behavior has been removed while preserving saf...