Bad 👎: Didn't think Strix Halo would drop in $ w/ Gorgon Halo ( see 🧵for RAM data )
Good 👍: 2nd Strix Halo on the way. Planned vLLM, Ray, RDMA RoCe, with k3s cluster on top for AI Inference and other fun stuff.
Wait till you see what I deploy in it.
Decision data 🧵
2-bit Muse Glimmer GGUF managed to call 100+ tools on just 14GB RAM. 🔥
Muse Glimmer did a complete repo bug hunt for 5 mins nonstop with: evidence, repro, fix, tests and a PR writeup.
Run and train it in Unsloth.
GitHub repo: https://t.co/aZWYAtakBP
The future will be dominated by two types of models
ULTRA SMART - large extremely intelligent 10T models that can do very complex work
ULTRA CHEAP - near free intelligence that can handle pretty much everything else
Today’s equivalents are Fable and Luna / Deepseek Flash
Offically on Pi coding agent now. I don't need "y" or "n" .
I `/edge-case` my well researched Spec > plug the holes > let it rip.
Sys prmpt tune by A/B testing diff prompts & checking code for new models ( see vid )
Meta’s Muse Glimmer is out. Open Source still cookin people!
30B open (Apache 2.0)
Unsloth says “strongest agentic model in its class”
Quant ~17-20GB, strong tool use, reasoning, recovery, multimodal.
Weights on Hugging Face, Ollama (MLX) + Unsloth/llama.cpp support live, full integrations rolling out. And Muse Spark 1.2 open weights coming soon per the Zuck (see photos)
Who’s testing it?
#Agents #LocalAI #OpenSourceAI
Introducing Muse Glimmer, an open-weight 30B-parameter model optimized for local, always-on agent workflows.
Muse Glimmer delivers strong performance on key agentic use cases and benchmarks compared with leading models in its size category, and is designed to run entirely on consumer hardware like a Mac or PCs with performant GPUs.
In keeping with our long tradition of sharing fundamental AI research, we’re releasing model weights under a permissive Apache 2.0 license.
🧵👇
I really like it. Esp even though I installed web search, it refuses to use it and instead pulls the package down and searches the code.
Come to find out this is a lot of people’s exp. What I noticed is it’s faster than checking the Docs. I have a lot of other processes around how I use it that I’ll publish soon. But loving it with Laguna S2.1
@HotAisle Exactly my line of thinking for what I’m setting up today w/ multi model bias.
Competing model biases to fight over how the code doesn’t live up to the criteria.
Great point for local too, route to lessor model for quicker edits.
Pi is so good. I was blown away by the fact that it doesn’t want to guide the model to searching the web and instead the harness will search the code base of the packages it pulls locally to figure things out.
Extremely fast and get things right the first time with my local Laguna S2.1 model
@Keldrik It’s actually engineering genius because instead of searching the web for structured data which totally messes up your prefix cashing and spec decoding it instead searches the package it downloaded locally. Pi is crazy simple and crazy good 👍
@rozario_ru83489 Gotta latch on too and interact with ppl in your niche then it starts getting good.
Reply to people in the comments you like their tech, whatever they are doing and your whole feed becomes that