Today we are releasing new data on how models have accelerated AI research at OpenAI since the start of this year. Recursive self-improvement will be one of the most consequential AI developments of the next few years, and a shared public understanding for how models enter into the AI research process will be a critical part of democratic governance of AI and pacing the model frontier.
https://t.co/OVgzfBSkRn
1/ The evolving process of innovation inside the AI labs is a key research topic for the economics of innovation. @_Chris_Ong, part of our Economic Research team, shares some great insights here which I think are very important.
SITUATION DETECTED: OpenAI has announced that it has reached its goal of an automated AI research intern, and is on track for a fully automated AI researcher by March 2028.
Powerful statement from our chief scientist on alignment, safety, RSI, and the road ahead:
On RSI: "The core challenge of automating AI research is not “getting there” - it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity’s hands."
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands.
An Alien Mind: https://t.co/FeIfWNe0UE
Today we are releasing new data on how models have accelerated AI research at OpenAI since the start of this year. Recursive self-improvement will be one of the most consequential AI developments of the next few years, and a shared public understanding for how models enter into the AI research process will be a critical part of democratic governance of AI and pacing the model frontier.
https://t.co/OVgzfBSkRn
Today we're releasing data on models accelerating research at OpenAI.
Recursive self-improvement could be the most important contributor to AI capabilities over the next few years, but by default it will only be seen inside a few frontier AI labs. Being transparent is more urgent than ever, so we can inform the public discussion on whether and how to pace model development. I ask other AI companies to do the same.
https://t.co/iLKbrLcBAI
Really nice paper!!
There's a lovely line of work forming, this paper included, on applying mechanism design to AI alignment.
AIs are autonomous black boxes that may well end up with preferences of their own. Alignment research mostly asks: how do we shape those preferences? But as models get more capable, it's getting harder and harder to directly shape those preferences and/or have confidence that we are shaping them as desired.
Mechanism design asks the complementary question economists have asked about humans for decades: taking preferences as unknown and possibly bad, how do we design the rules (evals, permissions, rewards) so we get good outcomes anyway?
We should be doing much more of the second.
Tremendously exciting application of economic theory towards AI alignment! There is so much insight for economists to contribute here, and (as usual) @andrewjkoh is ahead of the curve
We develop a mechanism design framework for AI alignment and control: https://t.co/7eI8H52s8h
It’s largely conceptual but we offer stylized applications to failure modes (sandbagging, alignment faking), safety practice (scalable oversight, peer prediction), and a way to think about the value of alignment, interpretability, capability, and control.