AI is getting cheaper more quickly than any other transformative tech in history. At a given level of performance, cost has fallen ~47%/quarter since 2023.
That’s 4× faster than DNA sequencing, 6× faster than compute, 18× faster than lithium batteries, and (up to 1973) 54× faster than electricity.
We can say we’ve reached #AGI once a model scores 100% on every benchmark, regardless of the version. The ultimate #AGI benchmark will be scientific discovery..
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
52.6% on Terminal-Bench-Science 0.1, > x2 Fable 5
long-horizon agents :
this shift to AI agents that can work autonomously for hours or days (coding, researching, running experiments, optimizing GPU, correcting their own work)
to me,this is what the singularity look like.
An idea I’ve been thinking about a lot lately.
I discussed it with #ChatGPT and started exploring the concept of a “machine economy” for the open web:
As AI agents start browsing the internet (ex: Grok bot...) at massive scale, traditional ads won’t work for machine traffic.
What if agents simply paid a tiny amount every time they consume information?
Micropayments → publishers → more incentive to keep creating and maintaining an open web.
We have paused some frontier RL training to ensure that we can meet the appropriate alignment, security and monitoring standards for the new level of capabilities in front of us. Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment.
We care very deeply about AI safety. We believe the entire field will have to coordinate on shared safety standards, but will act unilaterally in the meantime.
We expect confidence in safety to increasingly set the pace of AI progress. We are optimistic about the alignment work we are doing, and we remain committed to making frontier capabilities widely available.
https://t.co/51kvKfbfrO
Yesterday it was DS v4 and Grok 4.6.
Tonight we’ve got GLM 5.3 and Gemini 3.7 Flash.
The speed at which new models are dropping is honestly crazy. every night i woke up with new things!!!
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i
Releasing the model weights and technical report of Kimi K3.
Kimi K3 is our most capable model: a 2.8T MoE model with native visual understanding and a 1M-token context window.
New model architecture: 2.5x the intelligence per unit of compute, not just more params.
Alongside Kimi K3, we're opening up more of the stack behind it — high-performance attention kernels, MoE communication library, and infrastructure for running agent environments at scale.
Model weights: https://t.co/7m7eEg6Y0B
Tech report: https://t.co/yeu6cjpMCT
Tech blog: https://t.co/YTfiMSNM1f
Qwen3.8 is officially on the way.
2.4T parameters, open-weight, and the Max Preview is already available to test.
The open-model race just got even more interesting.
#AI#Qwenstronger.
#AI#LLM#OpenSource