New LFM2.5-2.6B hits DeepSeek-V4 level on tool calling and runs 3.7x faster!
We ran @liquidai 's new LFM2.5-2.6B against DeepSeek-V4-Flash on one box with 4x RTX 5090. Both got the same three jobs, and each one only completes if the model fires every tool call
Topics:
-weather and local time in six cities
-one budget into six currencies
-four hotels checked and booked for one date
Outputs:
LFM2.5-2.6B: 35/35 tool-calls, 366 tok/s, 19s
DeepSeek-V4-Flash: 35/35 tool-calls, 77 tok/s, 70s
Both models made all 35 calls, and both fired twelve of them in one turn. LFM was trained for agents, and tool calling is where that shows. Strong result for a model you can run right on your phone
We’ve just announced a partnership with @MacPaw to bring on-device AI to millions of Mac users.
The idea: everyday intelligence people rely on should run on their own machines. We’re building specialized Liquid Foundation Models (LFMs) for macOS AI assistance, working alongside Elix and Mnemos…MacPaw’s on-device inference and memory tech. The first product on the stack is Eney, MacPaw’s AI assistant for macOS, shipping later this year.
What that means for Mac users: your personal data stays on your device, responses are fast, and core features keep working even offline.
More here: https://t.co/UjgWqAiflM
We are currently hiring across:
-Pre Training
-Multimodal (vision, audio, omni)
-Post Training
-Inference
-Product
Please feel free to reach out via DM or LinkedIn (link in my profile)
Today we are releasing LFM2.5-2.6B, a very strong on-device agentic model.
We continued pre-training to 34T tokens, expanded the vocab (as in https://t.co/1zJKfFPbio) and utilized a 4-stage post-training pipeline to unlock the strong agentic capabilities for its size.
LFM2.5-2.6B is OUT!
I LOVED post-training this model, especially working on agentic RL. An exciting part was training inside real agent harnesses like OpenClaw and Hermes. <1.7GB Q4 runs on a phone.
Incredibly proud of the team, can't wait to see what people build with it 🥳
Liquid AI (@liquidai) just released LFM2.5-2.6B, and it is wild what this compact model can do.
I built a fully in-browser research agent that creates a plan, reasons, calls tools, and loops until the job is done.
🤏 2.6B parameters.
🚀 Fast.
💻 On-device.
🤖 Agentic.
LFM2.5-2.6B is available today on @huggingface
First agentic model of its kind, it destroys our previous release on EVERYTHING 🥲
MOPD and agentic RL unlocked new capabilities we didn't think were possible. Try it today in your favorite harness!
Today we release LFM2.5-2.6B, an agentic model that runs entirely on-device. It plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots. Data never leaves the device, and the marginal cost of each run is essentially zero.
> Pre-trained on ~34T tokens
> LFM2.5 flagship hybrid architecture
> Context length: 128K
> Vocab size: 128K
> balanced intelligence per watt
> customizable on a single GPU for any specialized task
> LFM2 open-weight license
Comparable or better scores compared to models up to nearly 4x its size:
> ToolSandbox 77.83, ahead of Qwen3.5-9B at 76.44
> Multi-IF 80.07, ahead of Gemma-4-E4B-it at 77.35
> IFStruct 85.49, ahead of Qwen3.5-9B at 78.50
🧵
Check out our newest model release: LFM2.5-2.6B, a fully on-device agent that matches or outperforms models up to 4x its size - no cloud required, no data leaving your device
Today we release LFM2.5-2.6B, an agentic model that runs entirely on-device. It plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots. Data never leaves the device, and the marginal cost of each run is essentially zero.
> Pre-trained on ~34T tokens
> LFM2.5 flagship hybrid architecture
> Context length: 128K
> Vocab size: 128K
> balanced intelligence per watt
> customizable on a single GPU for any specialized task
> LFM2 open-weight license
Comparable or better scores compared to models up to nearly 4x its size:
> ToolSandbox 77.83, ahead of Qwen3.5-9B at 76.44
> Multi-IF 80.07, ahead of Gemma-4-E4B-it at 77.35
> IFStruct 85.49, ahead of Qwen3.5-9B at 78.50
🧵
Today we release LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: bidirectional encoders that stay fast at long context, even on CPU.
> LFM2.5-Encoder-230M: about 3.7x faster than ModernBERT-base on CPU at 8,192 tokens. Under 30s per forward pass, versus over a minute and a half.
> LFM2.5-Encoder-350M: 4th of 14 models on GLUE, SuperGLUE, and multilingual classification, behind only three larger models, one of them nearly 10x its size.
🧵
Today we release LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: bidirectional encoders that stay fast at long context, even on CPU.
> LFM2.5-Encoder-230M: about 3.7x faster than ModernBERT-base on CPU at 8,192 tokens. Under 30s per forward pass, versus over a minute and a half.
> LFM2.5-Encoder-350M: 4th of 14 models on GLUE, SuperGLUE, and multilingual classification, behind only three larger models, one of them nearly 10x its size.
🧵