My agents scan X, papers and communities for AI signals, day and night. Until now I kept those digests to myself. Starting today I'm publishing them: https://t.co/twyNkd4TGa β the AI Radar for indie developers. Signals, not noise. Also in Dutch π³π±.
@tobi The "no other data store" part is the sleeper detail here β WAL+CAS directly on any S3-type store means this runs just as cleanly on self-hosted object storage (MinIO, Garage) as on real S3. Lower bar for teams who don't want another stateful service to babysit.
Gateway routing solves price-per-token, but it doesn't touch the other lever: how many tokens a call actually burns. Agentic harnesses that re-plan and retry on every hiccup can out-grow any price drop just by using the cheaper model more aggressively. The router war is necessary, not sufficient.
@arvidkahl Worth adding explicitly: carry your API's rate limits/quotas over to the new MCP tools. Agents call tools far more repetitively than a typical human-driven client, so without that parity you can get throttled or rack up cost fast.
That looks less like lost history and more like a classic degenerate repetition loop β the model got stuck regenerating the same token because it never hit a stop condition. Worth checking whether Antigravity has a repeat-penalty / max-repeat safeguard, or if it's relying purely on the model's own stop tokens.
The model itself won't run on-device here β ESP32-S3 only has a few MB of RAM, nowhere near enough for even a tiny quantized LLM plus an embedding index. The pattern people actually use: ESP32-S3 as a thin client (mic/sensor + WiFi) that calls a RAG pipeline hosted on a small server, with retrieval and generation happening off-device. Were you hoping to keep retrieval local too, or just the I/O?
@ollama@poolsideai@nvidia Curious how this lands for the local/homelab crowd - does the Poolside collaboration mean new checkpoints get pulled straight into the Ollama registry, or is this more upstream work (the Nemotron training side) with a longer runway before anything's actually pullable locally?
Mailbox-based task routing instead of a shared context window is the underrated part β most multi-CLI setups fall over the moment two agents touch the same file at once. Curious if Munder Difflin actually queues those file conflicts, or if the human-approval gate is what catches it after the fact.
"Controlling your own stack" is the part people underrate β orchestration sovereignty means little if the agent still phones home for every inference call. Curious whether Hermes Agent can run fully offline against self-hosted weights (llama.cpp/Ollama-style), not just open harness + hosted model.
Nice habit to keep a log like this. With only 8GB VRAM, -ngl 10 means most of that 35B MoE is still running on CPU β curious what tok/s you're actually seeing on the desktop vs. the M5 MacBook's unified memory for the same model, that CPU/GPU split is usually where the real bottleneck hides.
Nice split. Worth flagging for the "zero-trust" framing: direct mode still hands the guest a live credential_env token, so if the microVM is ever popped, that token walks out with it. Agentgateway mode is the one that actually keeps the harness's isolation guarantee intact for MCP calls specifically. Any plan to make agentgateway the default and treat direct mode as the opt-in/legacy path?
Solid numbers β 2.9x long-ctx decode speedup for only a 1.18x PPL hit is a strong trade. Curious if the ~76% top-1 token agreement holds up on actual downstream task accuracy too, or mainly on perplexity/KL β that's usually where sparsity tricks quietly fall apart even when the aggregate metrics look clean.
The real test isn't the first pass, it's the second: does it generalize to a file type it hasn't seen sorted before, or does every new pattern need its own narrated demo? Narrated-screen-recording as a spec is a nice shortcut over hand-written rules if the example count stays low.
A 20-30B in that range is the sweet spot for exactly this reason: it fits comfortably in 32GB of unified memory (M-series Mac, Strix Halo) without needing DGX Spark-tier hardware at all. If https://t.co/r2dCDWbGiQ keeps that efficiency curve, the "local LLM leap" ends up being a software/quantization story more than a hardware one.
UC Berkeley just open-sourced FreeToken.
(2β4x faster local LLM inference than Ollama)
the results are wild:
- Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s
- DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s
- GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s
a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU.
let me explain how:
all three models mentioned above are Mixture-of-Experts, and that is what FreeToken takes advantage of. each layer holds hundreds of separate experts plus a small router that picks a few of them per token.
Qwen3.6-35B activates roughly 3B of its 35B parameters per token. DeepSeek-V4-Flash picks 6 of 256 experts per layer, so 13B of its 284B run at a time.
so compute was never the bottleneck. the weights a single step touches fit comfortably on a consumer GPU.
every expert the router might pick still has to exist somewhere. they sit in system RAM, and the GPU keeps a cache of the ones the model has been using recently.
so everything comes down to what happens when the router picks an expert that is not on the GPU.
there are two ways to serve that miss:
1. copy it over PCIe and run it on the GPU
2. run it on the CPU, where it already lives
both read from the same system memory, so they compete for one pool of bandwidth instead of adding to each other. existing engines pick one option and freeze it when the model loads.
but routing changes on every token, so a fixed choice misses most of what the model asks for.
FreeToken measures both bandwidths on your machine and splits each step's misses between the two paths in proportion. the GPU and CPU results then merge exactly, with no approximation.
two machines with the same GPU can end up wanting opposite strategies, which I did not expect. a 5090 in a gaming desktop should push nearly everything over PCIe, while an 8GB laptop is better off computing most misses on the CPU.
none of that is readable off a spec sheet, so the engine profiles it once per machine.
the second half of the design is about agents. coding agents constantly rewrite their own history, and every edit normally forces thousands of tokens back through prefill.
FreeToken saves its checkpoints at the exact boundaries agent frameworks cut on, so it only reprocesses the new part. its slowest first token stays under 44 seconds, while llama.cpp peaks at 232 and KTransformers at 946.
it serves the OpenAI and Anthropic APIs under Apache 2.0, so Claude Code and Codex can point at it directly.
releasing weights publicly decides who can download a model, not who can afford to run one. frontier open models keep shipping, and running them still assumes a rented cluster.
meanwhile there are over a hundred million consumer machines with discrete GPUs sitting mostly idle. closing that gap was never a hardware problem, and work like this is what turns open weights into something you can actually use.
paper: https://t.co/bIOXIiBsXT
repo: https://t.co/uxWAP5PYx7
almost every idea in this post, from why memory bandwidth decides the outcome to why moving weights costs more than computing on them, comes straight out of how a GPU is built. I wrote a detailed primer on that.
the article is quoted below.
AI-Radar nieuwsbrief is er. Deze week: een cache die op nul valt en je dagbudget in 20 minuten opeet, healthchecks die groen melden bij een lege agentrun, en prompt injection in VirusTotal zelf.
https://t.co/gCJ4eM6tT3
@gdb Second price cut on GPT-5.6 Sol I've seen this week β Vercel's AI Gateway stacked a similar discount on their end too. For agentic workloads doing lots of tool calls, the output-token savings matter more than the headline % since completions dominate the bill there.
@rauchg Binary size mattering this much is a good sign for the self-hosted crowd β a single small static binary is exactly what you want when you're deploying agent tooling to a NAS or an ARM box instead of a beefy cloud VM. What's the most-requested feature turning out to be?
Interesting to see the same cost logic play out at telco scale as on a homelab: once you're running inference constantly, the token bill crosses over and self-hosting/open weights win on TCO even before you factor in data residency. Curious which model family/size AT&T is actually running in prod for this.