Liquid AI ships LFM2.5-2.6B — on-device agents with 128K context and open weights
Liquid AI’s LFM2.5-2.6B runs planning and tool-calling agents on phones and laptops, with ~34T pretraining tokens and open weights for on-device stacks.

On August 4, 2026, Liquid AI released LFM2.5-2.6B: a 2.6B on-device agent model with 128K context, about 34T pretraining tokens, and open weights that run in llama.cpp, MLX, vLLM, SGLang, and ONNX. The company posts 220 tokens/s on an M5 Max, 113 on a Ryzen AI Max+ 395, and about 30 tok/s on a phone, under 2.5 GB. Hugging Face has the Base and post-trained checkpoints.
This is not a tiny chat toy. Liquid trained it to plan, call tools, and live inside Hermes, OpenClaw, and Pi without a cloud bill.
Architecture and training stack
2.6B parameters, vocab doubled to 128K in place for non-Latin scripts, mid-training with a dedicated 128K context-extension phase. Post-training is a four-stage stack: two SFT rounds (the mix is about 7× the LFM2.5-8B-A1B SFT mix, heavier on tool use, search, SWE, and agent traces); teacher specialization with RLVR experts for instruction, math, knowledge, code, tools, and long context; Multi-Domain On-Policy Distillation so the student rolls out under its own policy while a routed teacher gives token-level feedback; then agentic RL in real harnesses with GRPO and a safety gate.
That last stage is the product. They trained inside the harnesses they want you to use.
Tool-use benchmark comparison
Liquid’s table has LFM2.5-2.6B beating or matching models up to about 4× larger on instruction-following and most tool-use sets. It leads every IF bench in the chart and nearly every tool bench, trailing Qwen3.5-9B on BFCLv4. Agentic scores beat both Gemma 4 E2B/E4B and trade with Qwen 3.5 4B/9B. STEM leads AA-Omniscience-Public; coding is where the bigger models keep an edge. Those are Liquid’s vLLM evals with published temperatures. Rerun if you procure.
On-device speed and memory
220 tok/s decode on M5 Max, 113 on Ryzen AI Max+ 395, 30 on a phone, <2.5 GB. On an H100, Liquid quotes almost 15K output tokens/s at high concurrency — about 1.3B tokens/day on one GPU. Day-one runtimes: GGUF, MLX, vLLM, SGLang, ONNX. The pitch is free inference, privacy, and massively parallel local agents because tokens no longer have a price.
Wiring Hermes/OpenClaw harnesses
Serve an OpenAI-compatible endpoint, point Hermes Agent, OpenClaw, or Pi at it. Liquid says it works out of the box. That matches the RL setup: the model already saw those system prompts and tools.
LFM2.5-2.6B is an August 4 open-weight agent for the edge. Download both checkpoints. If your workload is coding-heavy, Liquid already told you to pick a larger model.



