1/ big announcement today: we will be releasing an open weight version of muse spark 1.2 soon.
we also are releasing muse glimmer, a 30B agentic model with open weights under apache 2.0. muse glimmer can run on 24GB of VRAM without losing agentic reliability. 🧵
Great to see @AIatMeta back publishing open models 🙌
Muse Glimmer is a 30B open-weight dense model with a 120K+ context window, built for long-running agents, delivering up to 20K tokens/sec on a single GPU.
It’s optimized to run locally across NVIDIA edge, desktop, and workstation AI platforms.
Try it with our GPU-accelerated endpoint: https://t.co/pcyRAQWeyL
Benchmarks tell us whether a model can build something well.
But actual usage adds another layer: whether people find the result useful enough to keep interacting with it.
That makes this ranking interesting beyond Grok 4.5 taking the top spot. It starts connecting model capability with real product behavior.
As AI gets better at building complete apps, metrics like usage and retention may become just as important as measuring whether the app was generated correctly.
Grok 4.5 just ranked #1 on Design Arena for Daily Usage - measuring the average unique users actually interacting with apps built by each model every day
It's not just winning benchmarks....it's winning users
And it's outperforming Claude Fable 5, Opus 4.8 and GPT-5.6 Sol by a ridiculous margin:
1. Grok 4.5 — +563.2% vs average
2. GLM 5.2 — +241.3%
3. GPT-5.6 Sol — +103.0%
Claude Fable 5 — +39.2%
Claude Opus 4.8 — +38.9%
The interesting thing about this benchmark is that it's measuring whether REAL PEOPLE actually use the apps these models build
And Grok 4.5 is absolutely dominating
HUD mode
Hermes stops being a window you switch to and becomes a layer over the app you're both working in.
Or keep it around as a little buddy agent. Ask it random things, drag it anywhere, it's yours
@NousResearch
Every AI tool you need to escape the permanent underclass:
• Codex app (5.6 sol medium)
• Hermes Agent (powered by Qwen local model)
• OpenClaw (ChatGPT 5.6 oauth)
• Gemma 4 running on a Mac Mini
• ChatGPT Voice to get work done while getting steps in/getting lean af
• Claude Fable 5 for planning
• Claude Design for front end design
• ChatGPT Image gen 2 for more than you can imagine
• Spotify playing lofi bangers
• 2nd monitor that has these agents up 24/7
Do work on 1st monitor. Constantly prompt 2nd monitor
Personal finance could become one of the clearest use cases for AI agents.
Not because they can show you where your money went, but because they can find what you missed and help you act on it.
That’s a much more useful kind of intelligence.
Used ChatGPT finance just now and asked it to find phantom subscriptions. It unearthed $550 a year of random things I didn’t realize I was paying for and thought I canceled. Thanks @OpenAI and whoever is on the Finance product.
Claude Code can use your iPhone! 📲
Introducing: phone-harness
> Automate any iOS app like a human
> native iPhone control → no API, no jailbreak
> Connect once, control anytime
Setup in one prompt.
Try it now! ↓
GPT-5.5 will not be the last major pre-training run from OpenAI.
GPT-6 will be a great model. However, the end-of-year model I alluded to back in June is going to be OpenAI’s biggest pre-train, as far as I know.
Now we know that model is codenamed ‘Doug.’
And it will make Fable seem ‘primitive.’