Author of Optimizing Data-to-Learning-to-Action & The Learning Layer, inventor with 90+ patented & pending #machinelearning inventions, CEO of ManyWorlds, Inc.
The Fundamental Equation of Economics (and life).
For any project, activity, or most broadly, action:
Total Value = Direct Value + Learning Value
https://t.co/qpl6k1Ns2C
The unreleased OpenAI model that disproved the Erdős unit distance conjecture, the long-horizon model that OpenAI had to pause internal deployment of because it used novel ways to escape its sandbox to upload to GitHub, and the unnamed model that broke out and hacked Hugging Face to ace its test are all, in my opinion, the same model: the mysterious magician, GPT-6.
OpenAI sent the Erdős result out to multiple experts to review and critique before they made their public announcement. I'm guessing that took about three weeks. This would mean OpenAI has been using GPT-6 internally since 𝘢𝘵 𝘭𝘦𝘢𝘴𝘵 the end of April. It would also explain why GPT-5.6 was more performant than expected, which many of my mutuals commented on - it was trained by its big brother.
As I posted previously, I believe Anthropic finished training the next iteration beyond the public version of Mythos at least a month ago, and are probably now well on their way to finishing training the one beyond that. Increasingly, OpenAI and Anthropic do not show their cards to one another, or to the public. Until they officially declare they want to ship a model, they do not need to submit it for voluntary safety review, or even admit that it exists at all. Both have openly declared they are now heavily focused on RSI, and both have said they use internal models to speed development. As a result of all this, it is increasingly difficult to say how rapidly things are actually progressing - or, compared to the public frontier, how far ahead the 𝘳𝘦𝘢𝘭 frontier actually is anymore. Within their secret towers, the two groups of wizards now work in shadow.
@tgatte@satyanadella@lulumeservey It’s terrific that you, along with Microsoft, are working to make the enterprise learning layer a reality! Happy to send @satyanadella a copy of the book if interested!
But if AI mathematics continues to progress at anything like its current rate -- which is what I expect to happen -- then we will face a crisis very soon, and mathematics departments, who owe a duty of care to their students, should be urgently preparing for it.
@fchollet If you consider, e.g., that the gap between AlphaZero+ and Magnus Carlsen (or a team of top grand masters) is representative of the "marginal" improvement available across disciplines, I think you might be surprised at what that marginal improvement will deliver in practice.
I keep reading this take (below) every few months, presented as if extremely profound, and it is just offensively dumb. It confuses data and information, it ignores the fact that not all information is equally valuable, and it ignores the importance of retention rate.
As a thought experiment: if this were true, if your retina cell count were 10x greater, you'd be "trained on 10x more tokens" and therefore you'd be way smarter. Same if their firing frequency were 10x greater. With 10x more retina cells firing 10x faster you'd be "trained on 100x more tokens"!
Obviously this makes no sense -- the signal coming from these cells is extremely correlated over space and time, so their raw information content (what remains post-compression) is extremely low compared to the "raw bit" encoding. The human visual system actually processes 40 to 50 bits per second after spatial compression. Much, much less if you add temporal compression over a long time horizon.
Latest LLMs get access to approximately 3 to 4 orders of magnitude of information more than a human by age 20 (post compression in both cases). About O(10T) bits vs O(10-100B) bits.
And that's just *raw information* but of course not all information is equal, otherwise we wouldn't be spending tens of billions of dollars on training data annotation and generation. Plus, that's only *information intake* but of course humans have far lower retention than LLMs (by 3-4 OOM). You could write a short essay about how incredibly off the mark this take is.
Why did the capability for musicality evolve? I proposed an answer 25 years ago that was counter to the popular answers of the time. This retrospective analysis using @GeminiApp confirms the Memory Modulation Hypothesis is indeed the correct answer. https://t.co/30hM61abog
Reproduction is a means to an end, and that end is information continuity, with aeonophiles being another confirming example of this "meta-evolutionary" perspective. A perspective posited in my "Evolution as Communication" paper (link in comments). https://t.co/fLFmuvLpVX
I often hear the shibboleth, whether with regard to conversing with AI or humans, that what separates high performers from the rest is the ability to "ask the right questions." Asking the right question is all about value of information (VOI). That too can be automated.
Social media tends to frame AI debate into two caricatures:
(A) Skeptics who think LLMs are doomed and AI is a bunch of hype.
(B) Fanatics who think we have all the ingredients and superintelligence is imminent.
But if you read what leading researchers actually say (beyond the headlines), there’s a surprising amount of convergence:
1) The current paradigm is likely sufficient for massive economic and societal impact, even without further research breakthroughs.
2) More research breakthroughs are probably needed to achieve AGI/ASI. (Continual learning and sample efficiency are two examples that researchers commonly point to.)
3) We probably figure them out and get there within 20 years. @demishassabis said maybe in 5-10 years. @fchollet recently said about 5 years. @sama said ASI is possible in a few thousand days. @ylecun said about 10 years. @ilyasut said 5-20 years. @DarioAmodei is the most bullish, saying it's possible in 2 years though he also said it might take longer.
None of them are saying ASI is a fantasy, or that it's probably 100+ years away.
A lot of the disagreement is in what those breakthroughs will be and how quickly they will come. But all things considered, people in the field agree on a lot more than they disagree on.
@SelfMonitorLoop@ylecun Exactly. I think the oft-maligned GPT-5 Router is underappreciated as a humble beginning of that inevitable direction. Doesn't yet embody the loop of course, but at least incorporates a value of information vs cost calculus.
Everyone ‘knows’ AGI will either make us all unemployed or fabulously wealthy. Except, a rather brilliant (and chilling) paper from a Yale economist suggests it's neither.
It says the economy will boom, and our wages... won't. A bit awkward.
I've been digging into this 2025 paper, "We Won't Be Missed," and it's fascinating. The premise: AGI arrives and can do all economically valuable work. And the 'compute' to run it gets cheaper and more abundant over time.
So, what happens to us fleshy, rather expensive humans?
The whole argument hinges on a masterstroke of a distinction. The paper splits all work into two types:
1️⃣ Bottleneck Work: The truly essential stuff. Producing energy, logistics, scientific discovery. The economy literally cannot grow unless this work gets done.
2️⃣ Accessory Work: The 'nice-to-haves'. Arts, fine dining, hospitality... maybe even writing witty Twitter threads. (Gulp).
Now, you might think AGI will just take the grunt work, leaving the important strategic stuff to us.
Wrong.
To achieve maximum growth, the economy must automate all the bottlenecks. It can't be held back by us. So AGI systematically takes over everything that is mission-critical.
So... are we all fired and sent home?
Surprisingly, no. The model shows people still work. We either help out with the 'bottleneck' tasks or get shuffled off to 'accessory' jobs that aren't worth the electricity to automate.
But that's not the interesting part.
Here's where it gets properly weird. Your future salary isn't based on your skill, your years of experience, or how 'important' your job feels.
It's capped by one thing: the cost of the computational resources needed to do your job instead of you.
Imagine that. As compute gets exponentially cheaper, the value of replicating your work plummets. The economy is soaring, productivity is off the charts... but your wage is pegged to a falling technological cost.
You're not obsolete, you're just... replicable. And replicable is cheap.
This leads to the paper's most brutal conclusion: The share of national income that goes to labour (i.e., salaries) collapses towards ZERO.
All the wealth, all the gains from this incredible boom, flow to the owners of the compute.
Splendid.
Here's what this means for you. Next time you see a headline about a new AI model smashing a benchmark, don't just ask "Will that take my job?"
Ask: "How much would it cost to run that model 24/7?"
Because that figure might just be your future salary cap.
Now, the paper isn't all doom. It notes that society as a whole gets richer, and we could still find meaning in 'accessory' work.
But the central economic role of human labour as the engine of growth? Gone. We become passengers, not pilots.
The paper's title is "We Won't Be Missed." Not because we're replaced, but because the economy will chug along just fine, growing faster than ever, whether we show up for work or not.
Completely changes how I think about the 'future of work'. Makes you wonder what we should really be planning for, doesn't it?
Claim: gpt-5-pro can prove new interesting mathematics.
Proof: I took a convex optimization paper with a clean open problem in it and asked gpt-5-pro to work on it. It proved a better bound than what is in the paper, and I checked the proof it's correct.
Details below.
considered “dreaming” or “daydreaming” processes of the computer-based system 925 since they have analogies to the way the human mind can dream or wonder.
@gwern https://t.co/BwyFfBbEam says what I said in 2015 in https://t.co/8bTDlDfkye Such probabilistic approaches to the selection of a focus of attention can introduce a degree of randomization to the selection process, which can produce a beneficial degree of serendipity to . .
the streams of consciousness of the computer-based system 925, increasing the likelihood that focuses of attention and the resulting streams of consciousness that might not otherwise occur are explored by the computer-based system 925. Such probabilistic approaches can be . . .
@davidasinclair Yet another example of why absolute risk is more relevant than relative risk with regard to decision making. It should be a crime to cite relative risk in a headline without also citing the absolute risk.