Sulat.com
AI models
$10 off the fastest DeepSeek V4.1 Flash, Kimi K3 and GLM 5.3 from Synthetic
Together AI logo

Model details

LFM2-24B-A2B

LFM2-24B-A2B is a sparse Mixture of Experts language model that brings Liquid AI's hybrid architecture to its largest scale yet. The design pairs efficient gated convolution blocks with a small number of grouped query attention blocks—a combination discovered through hardware-in-the-loop architecture search to deliver fast prefill and decode at low memory cost. With 24 billion total parameters but only about 2.3 billion activating per forward pass, the model punches far above the weight of a typical dense 2B model at inference time. It was engineered to fit within 32 GB of RAM, making it deployable across cloud infrastructure, consumer laptops with integrated GPUs, and edge devices with dedicated NPUs. The LFM2 family has now scaled from 350 million to 24 billion parameters with consistent quality gains at each step.

The model builds on a training budget of roughly 17 trillion tokens using mixed BF16 and FP8 precision, and the architecture scales predictably—quality improves log-linearly across nearly two orders of magnitude. LFM2-24B-A2B is positioned as a fast inner-loop model for high-volume multi-agent pipelines, supporting function calling, web search, and structured outputs to handle multi-step workflows. With native support for nine languages and a 32K-token context window, it serves as a generation backbone in RAG pipelines and supports extended multi-turn conversations. Inference runs at around 350 tokens per second on an RTX 4090 and 293 tokens per second on an H100 GPU, with day-one support across llama.cpp, vLLM, and SGLang for flexible deployment.

Together AILiquidAI/LFM2-24B-A2Bliquid

Quick Info

Powered by
Provider
Together AI
Model key
LiquidAI/LFM2-24B-A2B
Release date
Feb 25, 2026
Last updated
Feb 25, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.03
Output token cost
$0.12

Limits

Output tokens
32,768 tokens
Context window
32,768 tokens

Transparent token rates

Compare LFM2-24B-A2B pricing

Rates are shown per one million tokens. Combined means one million input plus one million output tokens.

Browse this family

Latest news about LFM2-24B-A2B

Together AI

CoverageRelease Notes

The Opper AI release tracker lists LFM2 24B A2B among dated Liquid AI releases, placing its launch on February 25, 2026. The entry sits in a chronological index of ten Liquid AI model releases from September 2024 through August 2026, with the most recent being LFM2.5-2.6B on August 4, 2026. The tracker relies on Artifi For the specific LFM2-24B-A2B entry, the page exposes only a release date and an intelligence score field, with no architectural detail, benchmark breakdown, or changelog text in the supplied excerpt. As a third-party tracker rather than a primary Liquid AI source, the page is useful only to corroborate the model's rel

Videos about LFM2-24B-A2B

More models around LFM2-24B-A2B