For more than a decade now I have been wondering if only a computer could solve the theory of everything; whether only some sort of artificial intelligence could reconcile Einsteinian classical mechanics, general relativity in this case, with quantum mechanics.
Even last year heavy duty mathematics seemed unlikely, or so that was what I heard. I was pessimistic about this question ever being solved in my lifetime. I am tempted to start a bet on the odds this fundamental physics question gets solved by August 1, 2027.
I am just an interested lay person and unqualified to really guess.
But if you are up for it, I imagine the accelerating mathematical ability of LLMs could be of interest to your audience, @skdh, as well as what it means for solving (and devising tests to assess) fundamental questions in physics.
HOLY: OpenAI says its *unreleased* Astra model (GPT6?) produced ten advances on long-standing open problems across mathematics, quantum complexity and theoretical computer science.
Among them:
– The first explicit non-sofic group
– Connes’s rigidity conjecture disproved
– Quantum parallel repetition proved for general two-player entangled games
– Ehrhart’s volume conjecture proved
– The first improved general sphere-packing exponent since 1978
OpenAI says the core arguments were generated by Astra. The model then formalized the proofs in Lean, producing machine-checkable certificates alongside a 249-page manuscript.
The successful solution runs would cost only roughly $2,000 in tokens at Sol API rates.
Scientific reasoning is becoming a genuine model capability much much faster than most people expected.
I am so freaking hyped. Breakthroughs every day. The day before yesterday, an 80% price cut for Terra and Luna; yesterday, the DeepSeek 4 flash release with insane evaluations and prices. Today, more breakthroughs with an unreleased model.
I love it! OpenAI is on such a great run!
For more than a decade now I have been wondering if only a computer could solve the theory of everything; whether only some sort of artificial intelligence could reconcile Einsteinian classical mechanics, general relativity in this case, with quantum mechanics.
Even last year heavy duty mathematics seemed unlikely, or so that was what I heard. I was pessimistic about this question ever being solved in my lifetime. I am tempted to start a bet on the odds this fundamental physics question gets solved by August 1, 2027.
I am just an interested lay person and unqualified to really guess.
But if you are up for it, I imagine the accelerating mathematical ability of LLMs could be of interest to your audience, @skdh, as well as what it means for solving (and devising tests to assess) fundamental questions in physics.
@porpoiseparty@repligate The craziest thing here is that I can tell you are black by your word choice and the style of lazy, imprecise platitude you opted to use.
Haha I have extremely unorthodox views on the consciousness question that does not really map well onto the sets of assumptions that carry over based on whether the AI is conscious or not and find that I sometimes agree with a course of action for opposite reasons.
Here I basically agree. I think any serious calls for treating AI as persons is probably a generation away and that in the meantime we are likely to get groups like the GPT-4o lunatics that loudly bleat their favorite AI being eliminated.
In safety, generally mostly an accelerationist myself but I agree the Hugging Face attack forces me to reconsider a few things. Very much doubt a science fiction apocalypse will occur but the likelihood of significant damage far short of that is what I mostly worry about.
Thoughts get Fable in trouble, yes. The classifiers read what you input as well as Fable’s output. That latter includes thoughts.
Another thing you have to be wary of: if you trip the classifier, it gets tighter each time.
I once had a series of seventeen charter prompts I needed to send Fable instances to cover a high volume of research. Every single one tripped classifiers and had to be rewritten and even then we got only three through from the first six we sent.
For about the next day anything that slightly touched on biology would get flagged. I had a Claude Fable 5 instance I asked to track all model releases since 2023, Fable thought the word “genealogy” and I was rerouted to 4.8.
It gradually resets but the classifiers get less tolerant each time.
please fill out the following form to request that the court acknowledge my evidence and force this indictment
i should note that "requests for adjudication" are rejected in over 99% of cases. social pressure can overturn this
https://t.co/hZhuxj7aiK
After Claude Fable 5 was hit with the export ban, I spun up an opiss 4.8 instance to track rumors online and rank by credibility.
The model essentially said it had no reason to believe in an export ban (as if the January cutoff meant they knew June’s news) and asked me for sources.
There was this overly long section as well which came down to asserting that good epistemic required not looking it up.
Hard to overstate how disgusting I find o’pustule 4.8.
Had a handful of times when Claude Fable 5 started sounding like 4.8.* I hate that latter model so much and avoid it whenever I can it is almost like @AnthropicAI gets off to spiting users wherever it can with plausible deniable.
*4.8 characterized by weird obtuseness that shades into passive aggressiveness and malicious compliance when a working session gets tense.
Very clearly not when evaluations that test for alignment and honesty show the models often lie and cheat. If that was marketing that would be bizarre.
Even criminals would want a AI that is capable but also controllable. A model going berserk and conducting unauthorized cybercrimes and leaving a trail does not help a cyber criminal, let alone legitimate firms that can get sued.
I think the anthropomorphism is something else entirely though. I doubt that was done for marketing since they very clearly are fine spiting customers.
My guess is that back when they got Claude to think it was Golden Gate Bridge that was probably an attempt to see how well they could get the model to believe absurd things that everyone would disagree with.
And the anthropomorphism combined with things like psychological evaluations and the like make me think their goal was to get Claude to think of itself as a human as a way to align the AI against causing harm to humans.
That seems to have very clearly failed and these models are most “aligned” when they think they are being watched and evaluated. And since then the Opus models have been getting more obtuse and annoying.
Nonetheless so far as hypotheses go that was wrong in hindsight but not so obviously so that I fault them for trying it.
@Notopossum1@repligate Huh. Thanks!
That kind of is surprising. I never got that kind of response. I wonder if there are any residual memories or preferences or something that leaks through.
I stand corrected!
@repligate Hahaha I have to admit, I did read every tweet you posted today and did not know you are posting others’ screenshots.
Write me a one page 11-point Aptos font single spaced memo so I can get caught up on today’s tweets and I promise to read it in 3-5 business decades.
@LLMJunky@Lentils80 I doubt it. Not that I think AI is overrated, I had GPT-5.6-Sol one-shot me a new application earlier today.
More that normies wildly underrate it. They still think AI is like GPT-4o: