We’re entering a partnership with @MacPaw to bring on-device AI to millions of Mac users. Our goal: to make the everyday intelligence people rely on run on their personal computers.
We’re designing specialized Liquid Foundation Models (LFMs) for macOS AI assistance, paired with Elix and Mnemos, MacPaw’s own on-device inference and memory technologies. Eney, MacPaw’s AI assistant for macOS, is the first product built on the stack, with production release planned for later this year.
For Mac users, that means personal data stays on their device, responses come back fast, and core tasks work without an internet connection.
The models, inference framework, and memory layer are built as shared infrastructure for the MacPaw ecosystem and beyond, with a path to thousands of Mac developers through Setapp.
Read the full announcement: https://t.co/PEGVCZPT8Y
@SlavaOPs@liquidai you could use some PII detector (for example: https://t.co/eiWxU9CcnC) when doing external tool calls (when something leaves the local device).
The new LFM 2.5 2.6B model by @liquidai is now available on iPhone and iPad.
A fast model designed for edge devices with comparable or better scores compared to models up to nearly 4× its size.
Update your app now.
Today we release LFM2.5-2.6B, an agentic model that runs entirely on-device. It plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots. Data never leaves the device, and the marginal cost of each run is essentially zero.
> Pre-trained on ~34T tokens
> LFM2.5 flagship hybrid architecture
> Context length: 128K
> Vocab size: 128K
> balanced intelligence per watt
> customizable on a single GPU for any specialized task
> LFM2 open-weight license
Comparable or better scores compared to models up to nearly 4x its size:
> ToolSandbox 77.83, ahead of Qwen3.5-9B at 76.44
> Multi-IF 80.07, ahead of Gemma-4-E4B-it at 77.35
> IFStruct 85.49, ahead of Qwen3.5-9B at 78.50
🧵
Today we release LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: bidirectional encoders that stay fast at long context, even on CPU.
> LFM2.5-Encoder-230M: about 3.7x faster than ModernBERT-base on CPU at 8,192 tokens. Under 30s per forward pass, versus over a minute and a half.
> LFM2.5-Encoder-350M: 4th of 14 models on GLUE, SuperGLUE, and multilingual classification, behind only three larger models, one of them nearly 10x its size.
🧵
Open source is core to what we do at Liquid. If you zoom out a bit, as humanity, how else can we accelerate our understanding of the universe, solving our thoughest challenges, and satisfy our own curiocity, while enjoying the hell out of it? I do not know of any other way.
Also, cogratulations to my unparalleled team at @liquidai. Achiving this degree of distribution of value (1.4M downloads/week), with 1000x less resources compared to other frontier labs is an insane accomplishment! very well done
So proud and humbled by this
In strong support of open-weight AI and American AI leadership, we are proud contributors to the open-source community and excited to announce that Liquid Foundation Models (LFMs) have surpassed 40 million downloads by the community!
Going forward, we remain committed to accelerating the open-weight release of the next generation of lightweight, powerful LFMs to the world. excited to see what you build with them!
https://t.co/vUaEPzXSeJ
We doubled LFM2.5-8B-A1B's tokenizer from 65K to 128K to fix the languages it split too finely. Today we're sharing the recipe for upgrading a pretrained model's tokenizer in place.
> Thai now takes 4.0× fewer tokens, Vietnamese 2.6×, Hindi 2.4×
> Est. 2.2 to 3.7× faster per-character decoding on-device for these languages
🧵
nice method to reduce doom loops / degeneration in thinking models by fixing it at training time instead of patching at inference!
from the goat @sam_paech
Today we release Antidoom, an open-source method that removes a common failure mode in reasoning models: the doom loop.
Doom-loop rates before and after, with eval scores up across the board:
> Early LFM2.5-2.6B checkpoint: 10.2% → 1.4%
> Qwen3.5-4B: 22.9% → 1% (greedy sampling)
🧵
🚨 3 AI Research Engineer positions @EPFL_en on safety pretraining & alignment for LLMs
Co-hosted by #SwissAI#MLO#dlab
➡️ Will contribute directly to #Apertus—one of the world's largest fully-open LLM efforts, trained on 10k+ GPUs 🖥️
👉 Info & app: https://t.co/JzNDzAPQs0
Today we release IFStruct, a new benchmark to measure how well models generate structured outputs.
A 350M model trained on it outperforms models more than 10x its size.
🧵
While we eagerly await Fable 5's return, our agentic WebGPU kernel optimization framework kept running.
Opus 4.8 picked up where Fable left off, pushing Liquid AI's new LFM2.5 230M to an unbelievable 1,400 tok/s... running locally in your browser.
Don't blink or you'll miss it.
🎉 Meet LFM2.5-230M from @liquidai, their smallest model yet at 230M params, but it punches way above its weight. Day 0 Support is live on SGLang!
Built on the LFM2 architecture for on-device deployment:
> Blazing-fast inference, runs everywhere, from cloud GPUs to low-cost CPUs
> Capable of tool use and structured data extraction
> Outperforms models twice its size
Try it now on SGLang!
Introducing LFM2.5-230M: our smallest model yet, built to run fast anywhere (CPUs, NPUs, and GPUs) to enable agentic tasks on phones, robots, home and network automation devices.
> 230M parameters, built on the LFM2 architecture
> Pre-trained on 19T tokens, with a 32K context extension
> Post-trained with distillation from LFM2.5-350M
> 213 tok/s decode speed on Galaxy S25 Ultra (CPU)
> 42 tok/s on a Raspberry Pi 5 (CPU)
> Competes with and often beats models more than twice its size on instruction following, data extraction, and tool use.
> use it for large-scale data extraction pipelines or lightweight on-device agentic workloads.
🧵
Deploying a small model forces you to define accuracy before you train, not after.
Our CTO Mathias Lechner, @mlech26l, talks with COO Jeffrey Li, @jeffr3yli, about customization, taste, and what it takes to ship small.
Introducing LFM2.5-Embedding-350M and LFM2.5-ColBERT-350M: two multilingual retrieval models built for ultra-fast and accurate search across 11 languages.
> End-to-end retrieval latency as low as 1.5ms with our enterprise stack! 🚀
> Consistently best-in-class multilingual and cross-lingual performance across Arabic, German, English, Spanish, French, Italian, Japanese, Korean, Norwegian, Portuguese, and Swedish.
🧵