Using an "Agentic Ready OS" is changing the way I use computers…
Today I'm releasing the first alpha of the OmaPilot plugin for @OmarchyLinux and looking for more testers to iron out the Omarchy harness.
🤯 A 20GB model doesn't fit in this RTX 5080's 16GB VRAM. It still runs at ~100 tok/s.
🆕 NEW INFERENCE ENGINE!
FreeToken is a new open-source inference engine designed specifically to run huge MoE models on hardware that doesn't have enough VRAM to hold them.
... and there are new community tests to verify claims
🎮 RTX 5080 (w/16GB VRAM)
🧠 Ryzen 9 9950X3D
💾 64GB system RAM
🤖 Qwen3.6-35B-A3B NVFP4 (~20GB)
The model is larger than the GPU's VRAM.
Result:
🔥 ~100 tok/s
One example with a 1,028-token prompt reached roughly 🚀 110 tok/s
How?
FreeToken doesn't treat your PC as either GPU VRAM or system RAM. It treats the whole machine as an inference platform.
⚡ GPU-resident experts
🧠 CPU execution
💾 system RAM
🔄 dynamic expert caching
🚦 bandwidth-aware CPU/GPU scheduling
📚 dynamic KV-cache allocation
Because Qwen3.6-35B-A3B is an MoE, only a fraction of its total parameters are active for each token.
FreeToken exploits that sparsity instead of forcing the entire model into VRAM.
And the paper goes much further:
💻 8GB laptop GPU → 35B MoE
🎮 RTX 5090 → DeepSeek-V4-Flash 284B
🖥️ 96GB workstation GPU → GLM-5.2 753B
All on a single personal machine.
FreeToken is ...
🔓 Apache 2.0
🪟 Windows app
🐧 Linux
🖥️ GUI
🔌 OpenAI-compatible API
🤖 Claude Code / Codex / Hermes / OpenCode support
⚠️ The ~100 tok/s figure is a community result, not a benchmark reported in the FreeToken paper.
🔗 GitHub /FlashML-org/FreeToken
@NousResearch just opened Ox Alpha for free. A quadrillion tokens a day.
If you’re on Hermes, @Teknium already put `/model free` in. That’s the door.
Desktop app got 20% faster this week too.
Love your work team.
Yesterday, @blocks released a new open-source agent workspace called Berd. Unlike Buzz, this app's focus is on the solo builder working with multiple agent harnesses (who likes this funky style)
I made a quick demo showing how to install it, start talking to specialized agents, and even make your own agent based on your needs (mine being cat-themed visualizations). Check it out!
call me insane... but grok bot is the single most powerful agent stack i've ever used.
it will have your entire business run by agent teams working for you 24/7.
I just cancelled chatgpt and hermes to replace them with it
then mapped the entire grok bot agentic workflows into 4 excalidraw diagrams... (steal this)
→support desk: a chief of staff bot routes tickets, drafts replies, pings me only for refunds or anger
→money + ops: a bot processes invoices in gmail and keeps the books moving without me
→content: a bot scopes the brief, cites sources, hands me a publish-ready draft
→build: a bot reproduces the bug, files the ticket, hands the fix to the next bot
plus it's all free..give the screenshots to grok bot and it runs your business
New in Hermes Agent:
Have Hermes do an operation or set of operations on a website, and it can watch the api calls made there - then can create a static api for your agent or scripts it builds to use forevermore with this new optional skill!
Just run:
`hermes skills install official/web-development/har-derived-api-client`
Bookmarks saved on X pile up like a mountain, opened once and never looked at again
There's an open-source AI tool 'Siftly' that can solve this problem
2.5K Stars, it can really find those messily saved tools and articles
Import bookmarks → AI reads full text and screenshot text → automatic summary and categorization → natural language search + mind map browsing
Runs locally, no cloud upload, supports exporting CSV/JSON/ZIP. Using Claude saves you the API key hassle
Project address::
https://t.co/cymTfsIr5Y
Hermes Agent now runs Buzz.
The self-hostable workspace from @blocks puts humans and agents in the same messaging channels and codebase.
Three ways to use Buzz with Hermes (and vice versa):
- Buzz Desktop auto-discovers your Hermes install runs it locally
- A relay bridge gives it a hosted identity in your channels
- Connect via the Hermes Gateway to use Buzz as a full external platform with channels, DMs, threads, reactions, and cron delivery
https://t.co/srljIbERN7