GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
In fact, the continuous harness version significantly outperforms our human baseline in action efficiency across almost all levels. When we examined the reasoning chains to understand how the model operates, we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level. It goes as far as developing its own shorthand DSL to represent in-game situations -- essentially a game-specific algebraic notation.
Overall, Astra exhibits symbolic modeling behaviors we had previously only seen with sophisticated harnesses -- so harness capabilities are increasingly shifting into the model itself.
We see Astra as a major breakthrough in model intelligence.
Read our post on Astra and what these results mean: https://t.co/wJnYxEqYNI
After @OpenAI Astra’s release yesterday scoring 99.9% on the ARC-AGI-3 benchmark, I think it’s become clear that AGI is no longer some distant hypothetical. About 15 minutes after the benchmarks came out, I booked the first flight I could find to Jackson Hole.
I’ve now decided to relocate here permanently and begin establishing a life with as little exposure to AI as practically possible.
I’m currently evaluating houses based primarily on distance from town, whether they have a wood-burning stove, whether the water comes from a well, and how many important systems can still function if the internet disappears forever.
I’ve also started transitioning my daily life away from anything that could plausibly become agentic.
I paid for coffee with cash this morning. I asked a human being for directions. I bought a paper map. I have begun referring to Google Maps as “the old system.”
I’m also looking into purchasing a horse. At this point horses remain one of the only major transportation platforms with no API, no OTA updates, and no realistic path to MCP support.
I’ve started moving important documents to paper. I bought a typewriter and I’m learning which vegetables can survive a Wyoming winter.
Longer term, I’m hoping to assemble a small community of people with complementary pre-AGI skills: farming, carpentry, dentistry, medicine and cooking. LMK if interested.
Whatever happens next, I’ve made my choice. I’ll be in Wyoming splitting firewood and waiting for the compute to run out.
Today we're also opening the weights for Muse Glimmer, a great 30B parameter dense model that can run locally. Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model. Meta is a strong supporter of open source and I'm proud of these releases. Congrats to @alexandr_wang and the MSL team for all your great work on these models.
We pointed Muse Spark 1.2 at a kernel optimization task and let it run. 1,000+ tool calls over 24 hours on NVIDIA Hopper. It kept finding substantial improvements well beyond the initial exploration phase.
Releasing Muse Code in beta today. It's a terminal coding agent that takes on complete software engineering tasks across large repos: planning changes, writing code, validating the results. Powered by Muse Spark 1.2, a coding-focused model update.
Announcing Discovery Loop!
I am very excited to announce that, along with my longtime friends and collaborators @Sanjay_Ghemawat, @OriolVinyalsML and @quocleix, we are founding Discovery Loop (@DiscoLoopAI), a Public Benefit Corporation whose mission is to automate machine learning, science, and engineering to accelerate discoveries and progress. The four of us have worked together for 14 to 30 years, and have helped build some of the world’s most used products, infrastructure and AI models, and we’re excited to turn our attention to this ambitious endeavor.
♾
Learn more at: https://t.co/Rv3LMdLluK
@JensenHuang@nvidia I too am a big fan of open weight models, but open source AI doesn't stop there. Push for an overall "architecture of participation." Open source and the web won because they made it easy for people to cooperatively build out the system. Modularity and open protocols matter too!
Don't miss the chance to hear from Federico Castanedo, Director of Applied Science at Inception, as he explores the transformative power of AI in business value creation GITEX GLOBAL !
📅 16 October, 14:00
📍 Location: Hall 6, Stand B40 (G42 Booth)
The countdown to GITEX Global 2024 is on! We’ll be at the G42 booth, H6-B40 at the Dubai World Trade Centre from October 14-18, 2024. Don’t miss our interactive demos and expert sessions!
📅 3 more days to experience the future of AI!
Learn more:https://t.co/WUUgSXfkVk
> independently discover a Zeno's paradox at age 3
> MIT at 17, grad level math in 1st year
> graduate in 3 years
> drive motor scooters from Boston to Bogotá with the boys
> start a company in Colombia
> start code breaking with the IDA for money
> solve minimal varieties in riemannian manifolds
> speak out against Vietnam War, get fired from IDA
> take over math dept. at Stonybrook, make it a top-ranked program globally
> develop Churn-Simons theory, accidentally contribute more to physics than most physicists
> get bored with math, start modeling financial markets
> return 60% for 4 decades straight
> establish one of the most effective philanthropic organizations of all time
> chain smoke cigarettes the entire time
RIP Jim 🫡
Today, we are releasing Stable Video Diffusion, our first foundation model for generative AI video based on the image model, @StableDiffusion. As part of this research preview, the code, weights, and research paper are now available.
Additionally, today you can sign up for our waitlist to access a new upcoming web experience featuring a Text-To-Video interface.
To access the model & sign up for our waitlist, visit our website here: https://t.co/IcuPJr45S9
My full interview with @satyanadella about the state of play, why @sama actually got fired, and who will be the CEO of @OpenAI tomorrow.
https://t.co/j6NDnjwzJk
One of the remaining 5 employees at OpenAI should
- Opensource GPT4/Dall-E 3 weights
- Opensource repo
- Release GPT(n) training checkoints
- Publish training data
- Release all internal documentation
While they still have access