LFM2-24B-A2B is a sparse Mixture of Experts language model that brings Liquid AI's hybrid architecture to its largest scale yet. The design pairs efficient gated convolution blocks with a small number of grouped query attention blocks—a combination discovered through hardware-in-the-loop architecture search to deliver fast prefill and decode at low memory cost. With 24 billion total parameters but only about 2.3 billion activating per forward pass, the model punches far above the weight of a typical dense 2B model at inference time. It was engineered to fit within 32 GB of RAM, making it deployable across cloud infrastructure, consumer laptops with integrated GPUs, and edge devices with dedicated NPUs. The LFM2 family has now scaled from 350 million to 24 billion parameters with consistent quality gains at each step.
The model builds on a training budget of roughly 17 trillion tokens using mixed BF16 and FP8 precision, and the architecture scales predictably—quality improves log-linearly across nearly two orders of magnitude. LFM2-24B-A2B is positioned as a fast inner-loop model for high-volume multi-agent pipelines, supporting function calling, web search, and structured outputs to handle multi-step workflows. With native support for nine languages and a 32K-token context window, it serves as a generation backbone in RAG pipelines and supports extended multi-turn conversations. Inference runs at around 350 tokens per second on an RTX 4090 and 293 tokens per second on an H100 GPU, with day-one support across llama.cpp, vLLM, and SGLang for flexible deployment.