Sulat.com
AI models
NanoGPT logo

Model details

LFM2.5 2.6B

Designed for on-device agentic workloads, this 2.6B-parameter dense model sits within a family of hybrid architectures optimized to run locally while still handling long, multi-step tasks. Its backbone combines 22 double-gated short convolution blocks with 8 grouped-query attention layers across 30 total layers, and it was pre-trained on a 34-trillion-token budget before being post-trained specifically for agentic behavior. The result is a text-only model with native tool calling that has been exercised inside popular agent harnesses such as Hermes Agent, OpenClaw, and Pi, so it slots into existing orchestration flows with minimal plumbing.

In practical terms, the model is small enough to run on a laptop or phone yet ambitious enough to compete with substantially larger systems on tool use, instruction following, and multi-step reasoning. The team reports 220 tokens per second on an Apple M5 Max and 113 tokens per second on an AMD Ryzen CPU, all within a memory footprint under 2.5 GB, with the cataloged API limit context window that comfortably fits long tool traces and agentic scratchpads. Open weights ship across Hugging Face Transformers, GGUF, MLX, and ONNX checkpoints, and inference is supported through llama.cpp, vLLM, and SGLang, making it a flexible fit for local research assistants, edge-deployed copilots, and privacy-sensitive enterprise tooling.

NanoGPTliquid/lfm-2.5-2.6bliquid

Quick Info

Powered by
Provider
NanoGPT
Model key
liquid/lfm-2.5-2.6b
Release date
Aug 12, 2026
Last updated
Aug 12, 2026
Input modalities
Output modalities
Capabilities

Cost

Input token cost
$0.10
Output token cost
$0.20

Limits

Input tokens
128,000 tokens
Output tokens
32,768 tokens
Context window
128,000 tokens

Latest news about LFM2.5 2.6B

Videos about LFM2.5 2.6B

Recent tweets and retweets from NanoGPT

More models around LFM2.5 2.6B