Currently listed through these providers:
Model details
LFM2.5 2.6B
Designed for on-device agentic workloads, this 2.6B-parameter dense model sits within a family of hybrid architectures optimized to run locally while still handling long, multi-step tasks. Its backbone combines 22 double-gated short convolution blocks with 8 grouped-query attention layers across 30 total layers, and it was pre-trained on a 34-trillion-token budget before being post-trained specifically for agentic behavior. The result is a text-only model with native tool calling that has been exercised inside popular agent harnesses such as Hermes Agent, OpenClaw, and Pi, so it slots into existing orchestration flows with minimal plumbing.
In practical terms, the model is small enough to run on a laptop or phone yet ambitious enough to compete with substantially larger systems on tool use, instruction following, and multi-step reasoning. The team reports 220 tokens per second on an Apple M5 Max and 113 tokens per second on an AMD Ryzen CPU, all within a memory footprint under 2.5 GB, with the cataloged API limit context window that comfortably fits long tool traces and agentic scratchpads. Open weights ship across Hugging Face Transformers, GGUF, MLX, and ONNX checkpoints, and inference is supported through llama.cpp, vLLM, and SGLang, making it a flexible fit for local research assistants, edge-deployed copilots, and privacy-sensitive enterprise tooling.
Quick Info
Powered by- Provider
- NanoGPT
- Model key
- liquid/lfm-2.5-2.6b
- Release date
- Aug 12, 2026
- Last updated
- Aug 12, 2026
- Input modalities
- Output modalities
- Capabilities
Cost
- Input token cost
- $0.10
- Output token cost
- $0.20
Limits
- Input tokens
- 128,000 tokens
- Output tokens
- 32,768 tokens
- Context window
- 128,000 tokens