Liquid AI Releases LFM2.5 For On Device AI Agents
Liquid AI has released LFM2.5-2.6B, a 2.6 billion parameter model built to run capable agents entirely on the device. It supports tool calling and multi-step workflows while staying small enough for everyday hardware, from laptops to phones. The pitch to developers is that agents can be deployed anywhere, data stays private on the device, and usage scales without a cloud inference bill.
Transcript
Liquid AI released a 2.6 billion parameter model today that runs capable agents on a phone, no cloud needed.
LFM2.5 powers capable agents entirely on device. It supports tool calling and multi step workflows. Developers keep data private on the device and avoid a cloud inference bill.
Liquid AI says it competes with models 4x larger on tool use, instruction following, and agentic tasks. An Apple M5 Max runs it at 220 tokens per second.
It needs under 2.5 GB of memory. Both models are on Hugging Face today, with day one support for llama.cpp and MLX. Coding is where larger models perform better.
Inboxsmith helps small businesses handle calls and messages so nothing gets missed. Please like and subscribe for more news.
Sources
Every claim in this video comes from the top ranking coverage of this topic. The claims and where each one came from:
- LFM2.5-2.6B is built to power capable agents entirely on-device.(Liquid AI's official announcement)
- At 30 tokens/s, it allows you to run capable agents even on a phone.(Liquid AI's official announcement)
- It supports tool calling and multi-step workflows while staying small and fast enough for everyday hardware, from laptops to phones.(Liquid AI's official announcement)
- This enables developers to deploy agents everywhere, keep data private on the device, and scale usage without a cloud inference bill.(Liquid AI's official announcement)
- Best-in-class agent: Competitive with models 4x larger on tool use, instruction following, and multi-step agentic tasks.(Liquid AI's official announcement)
- LFM2.5-2.6B is the fastest model we tested, with decode speeds of 220 tokens/s on an M5 Max and 113 tokens/s on a Ryzen AI Max+ 395.(Liquid AI's official announcement)
- Efficient inference: 220 tok/s on an Apple M5 Max and 113 tok/s on an AMD Ryzen CPU, in under 2.5 GB of memory.(Liquid AI's official announcement)
- Both LFM2.5-2.6B and LFM2.5-2.6B-Base are available on Hugging Face today.(Liquid AI's official announcement)
- LFM2.5-2.6B ships with day-one support across the inference ecosystem, including llama.cpp, MLX, vLLM, SGLang, and ONNX.(Liquid AI's official announcement)
- Coding is the one place the larger models keep a clear lead, so reach for something bigger there.(Liquid AI's official announcement)
