Australia has been hacked.
'And today, I spoke with the CEO of OpenAI, Sam Altman, to express Australia's extreme concern about this incident. And I also expressed my disappointment that it took the company way too long to inform the government what had occurred, and the nature of the way that that notification occurred as well was unacceptable.'
“You AI people are so naive, I live in the REAL world, where [I have been consistently wrong in my predictions about AI and behind the ball on every trend related to AI, routinely making little prognostications that stochastic gradient descent joyfully stomps all over.]”
@MasterTimBlais i imagine that it similar to asking someone that speaks mandarin how many strokes it takes to make a certain character/word. or even english really.
My takeaways from today: OpenAI did a biblically-awesome thing (proposing a solution to Navier-Stokes!) with less-than-pure motives to say the least (beating an independent duo loosely affiliated with Anthropic to the punch) while probably not violating serious side constraints (deliberately cheating) while seemingly violating some minor ones (putting weird pressure on the duo in their offer of authorship) -- and plausibly, but perhaps not likely, benefiting from training on the usage data of the duo, who presumably had the disadvantage of working slowly over a year and not dropping however-many-millions on it over the course of a single week.
I don't know how the mathematics community will or should respond, but I do find Terry Tao's take a bit concerning:
@ruben_bloom In your estimation, would you say that the difference you are experiencing (going from net neutral/negative to 10x) is largely because wrangling those old models to do anything productive was not worth it compared to the effortlessness of models today?
GPT-6 Astra appears to be a massive jump in opaque reasoning ability: it looks like it can solve hard competition math problems entirely in its head (as in, without verbalized reasoning) while prior AIs could solve basic word problems.
This seems extremely concerning!
That is, if these benchmark results are representative (see the highlighted caveats in the image, I'm particularly worried about contamination).
Related to this, UK AISI found Astra has much worse monitorability.
I'd guess this jump is downstream of architectural changes (with increased serial depth) though a normal large pretrain scale up is a plausible cause. If the next few model generations involve similar jumps (presumably these jumps would be downstream of a transition to full-on opaque reasoning architectures with extreme depth), then chain-of-thought would no longer be a meaningful oversight tool.