AI productivity tools shouldn’t make you more productive at managing your AI productivity tool.
If I need prompts, agents, workflows, folders and a weekly maintenance ritual, we’ve somehow recreated the problem we were trying to solve 🥲
@liquidai’s LFM2.5-2.6B running agentic workflows entirely on-device is exactly the direction I want more models to go.
2.6B parameters. Tool calling. Multi-step tasks. Your laptop.
Small models getting better at actually doing things is much more interesting to me than another gigantic benchmark winner.
I think the next step after AI that can answer “what did I do today?” is AI that can say:
“You said you’d send this three hours ago and I haven’t seen you do it yet.”
That’s when computer memory starts becoming genuinely useful.
Thoughts?
Hey Prince! MacBook Pro M5, 24GB. Model: LiquidAI/LFM2.5-2.6B-MLX.
I dug into it and found my mistake: LM Studio had pulled bf16 for MLX, while GGUF was running Q5, so the comparison wasn’t fair.
Re-ran both matched at 5-bit (mlx-lm 0.31.3, mlx 0.32.0, pp512/tg128):
MLX: 57.4 tok/s gen, 1795 tok/s prompt
llama.cpp b10310: 56.6 tok/s gen, 1516 tok/s prompt
So basically identical generation speed, with MLX ~18% faster on prefill. My “MLX is slower” result was entirely down to the quant mismatch.
Happy to be wrong on this one lol sorry for panicking @Prince_Canuma & @ivanfioravanti
@ivanfioravanti@liquidai Yep! We integrated it into EverPaper. It can even message you through Hermes during the day to remind you about things it picked up from your screen or meetings. Only downside so far is speed: ~25–30 tok/s with MLX vs ~60 with GGUF.
What if your design handoff was just… pictures?
No Figma inspect, no redlines. Five rendered states + a motion table in a README. The build agents read the actual images; the reviewer diffs the code against the pictures, and rejected the first pass because a paused ring wore the wrong color.
Renders as binding spec. It works better than any document I've written.
What's your workflow for UI/UX Design?
(screenshots from @getEverPaper )
While developing @getEverPaper , I kept running into the same problem: coding agents can show you a perfect plan and still have permission to touch the entire project.
So I built TaskFence. I now use it to turn every approved plan into an enforceable contract: exact files, exact commands, automatic checkpointing, fail-closed enforcement, receipts, and rollback.
Open source (and FREE!):
https://t.co/LvUjYinoZT
@jacobAlswin Working on @geteverpaper (https://t.co/nevpzofQbB) an app that keeps a quiet daybook of your work.
Your screen, your meetings, your decisions. Everything is written down on-device.
Nothing leaves your Mac unless you say so.
Alright, @liquidai’s new LFM2.5-2.6B Agentic model is seriously impressive on @getEverPaper . Tool calling is extremely precise, and so far it’s been noticeably better than any other consumer-ready model I’ve tested.
It’s already available to EverPaper beta testers. Not in the beta yet? Comment here, or try it with your favorite agent harness.
This is just crazy!
I can’t wait to implement it in EverPaper. Everyone on X seems to have a DGX Spark or an overpowered laptop, but imagine how this could change everyday work for people on standard consumer hardware.
On it!
Today we release LFM2.5-2.6B, an agentic model that runs entirely on-device. It plans, calls tools, and works through multi-step tasks on phones, laptops, PCs, and robots. Data never leaves the device, and the marginal cost of each run is essentially zero.
> Pre-trained on ~34T tokens
> LFM2.5 flagship hybrid architecture
> Context length: 128K
> Vocab size: 128K
> balanced intelligence per watt
> customizable on a single GPU for any specialized task
> LFM2 open-weight license
Comparable or better scores compared to models up to nearly 4x its size:
> ToolSandbox 77.83, ahead of Qwen3.5-9B at 76.44
> Multi-IF 80.07, ahead of Gemma-4-E4B-it at 77.35
> IFStruct 85.49, ahead of Qwen3.5-9B at 78.50
🧵
we’re open sourcing a voice model built for how people actually speak.
20+ languages.
native accents.
mid-sentence language switching.
Expressive
yours soon.