Laguna XS.2 is a text-to-text model from poolside built around a Mixture-of-Experts design with roughly 33B total parameters but only about 3B activated per token, a sparsity profile that keeps compute and memory demands in check while still supporting larger-capacity reasoning paths. The architecture mixes Sliding Window Attention with a global attention layout, applying sigmoid-gated per-head routing in 30 of the 40 transformer layers to limit KV cache growth and speed up inference on a single workstation. Poolside developed it for agentic coding and other long-horizon work that benefits from running locally, and the weights are openly published so developers can self-host the model with quantization or run it through a hosted endpoint.
The model targets practical developer workflows rather than broad multimodal coverage: it accepts and returns text only, supports tool calling for agent loops, and ships with guardrail controls for safer integration into pipelines. A third-party feature listing reports a context input size around 131.1K tokens, which gives it room to hold sizeable code repositories, multi-file refactors, or extended task histories in a single prompt. Its combination of small active-parameter footprint, open weights, and agent-oriented capabilities makes it a reasonable fit for teams who want to run coding assistants on their own hardware while still being able to orchestrate tools and longer-running tasks without offloading everything to a remote service.