🚨 BREAKING: Microsoft confirms on publicly accessible web page that OpenAI has been using Looped Transformers in the GPT-6 series, proving The Information's reporting was correct all along‼️
GPT-6.1 Sol uses 2 inference passes, with a passing mention of "instead of three" 👀
Congrats my Reflection friends with this model release. Benchmarks are not the only thing that matters. If you can demonstrate a strong foundation of building the entire pre/mid/post training stack, then you are in the business. A few more rounds of iterations will get you there.
Introducing Beam: a highly efficient agentic open model with 501B total parameters and 23B active.
- Frontier reasoning efficiency
- Advances the Western open frontier on coding & agentic tasks
- Trained end-to-end from scratch
Full weights release this month.
Learn more about Beam: https://t.co/c3Qx2cpM8G
Can LMs discover & fix bugs if you don't tell them what went wrong? To proactively maintain a repo, agents need to find problems before users do.
SWE-sweep benchmarks this on 100 repos, 22 languages, 4k real bugs (numpy, php interpreter, lean kernel). Top models get <5%.
When building @Muse Code, the team spent a lot of time testing and hardening security, iterating for many weeks on guardrails and built-in skills to make the agent safer to use. It's great to see our robust security measures validated by this paper.
"We demonstrate that local agents in popular harnesses such as Claude Code, Codex, OpenCode, Antigravity, Grok Build, Muse Code, ZCode and Kimi Code can in fact tamper and delete entire traces when deployed in full-access mode across all three settings, with Muse Code being significantly safer."
(Holds true even on --yolo mode!)
Full text: https://t.co/OFuJAKqKJv
$META
Muse with a gangbusters start. Good job @finkd@alexandr_wang
Sensor Tower data:
Muse US downloads set a new record on 9/19 (264K), 3rd straight day >200K, while DAUs accel'd to a new peak of 448K on 9/18 (10-days post launch).
For context, ChatGPT didn't reach >200K US downloads until 1-yr post launch and 450K US DAUs until day 49. In the last 30-days, Muse has the highest avg rating (4.66 stars) w/ 87% 5-star review share vs peer apps.
Wanted to share my third AGI moment, after ChatGPT and Claude Code.
ChatGPT planned my trips, but it never booked them. This week @Muse booked every restaurant for our Japan trip, then went back and forth in Japanese with our Hakone hotel to secure a dinner slot that's not even on their website. The model generalizes across different tools so elegantly, to a degree that you would start to wonder if the actions were done by real humans.
Muse for Mac is out today! It works across apps, files, calendar, notes, and messages on your computer. You control what it can access. The team is shipping fast. Download at https://t.co/BX3oX19Zcl
It’s counterintuitive given Meta’s size, but Meta Superintelligence Labs feels more like a startup than any other major lab.
The teams building the models and products are shockingly small, and people have an unusual amount of agency. It’s a big part of why people choose to work here.
the most underrated part of muse is the model that powers it. thank you to the MSL research team—some of the most technically brilliant, hardworking scientists and engineers i’ve ever known. you brought original ideas to difficult problems, invented new methods, and battled through enormous technical complexity to build the jewel at the heart of muse.
thank you to the teams working on agentic post-training, tool and computer use, memory, personalization, alignment, safety, and multimodal understanding. thank you to the people building data, evals, and infrastructure beneath it all. so much scientific ingenuity lives in those details: designing a better experiment, finding the flaw in an apparently good result, staying with a stubborn problem until a new way forward becomes clear.
and now that jewel powers a beautiful product your grandma can use without knowing a thing about the extraordinary complexity underneath. it will just work for her. all those discoveries, all that patient work, becoming something simple and useful in another person’s hands. there is something magical about that. i’m grateful to be building with you.
Musecases just got 12 new real use case cards from X.
Wild ones this round:
- AT&T fiber bill negotiation (@chanduiiit)
- IKEA return end-to-end (@armand_ruiz)
- doctor admin in ~5 minutes (@armand_ruiz)
101 use cases now.
https://t.co/sKcZe3gnpO