The success of ChatGPT shows the power of Reinforcement Learning from Human Feedback. In this paradigm, models are aligned to values using a "reward model" trained on human preferences. I argue that it is imperative these reward models be open source:
https://t.co/sIk88viRYr
1/
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
@OpenAI It does increasing seem like now could be a reasonable time to pause additional scale-ups of pre-training as the world works through what gains can be made in efficiency, harness design and post-training
GPT-6 Astra is pretty insane at Minecraft! It was able to piece together and debug a full microprocessor + 8x8 pixel display from searching the internet, debug end-to-end by writing its own mod and test cases and record and edit a video for it end to end
At the core of our mission is working through how to ensure increasingly powerful AI benefits everyone.
We believe that, at some point in the future, AI acceleration for frontier model development may be so high that the world will need to pace the rate of AI advancement.
We hope to contribute to work led by the U.S. government, alongside other labs and the open-source community, to develop the tools and mechanisms that could make that possible.
https://t.co/pMCtiQjMoo
link to interactive tier list: https://t.co/AZcHgD6D6K
Disclaimer: this is just claude's assessment, but at least S-tier and some of A-tier look reasonable to me
For those following the recent vibe-math progress: I had Claude pull together an interactive tier list of 155 famous problems across 12 branches of math including Millennium, Hilbert, and Erdős problems. (Jacobian Conjecture: middle of A-tier. All 7 Millennium problems: S.)
@OwainEvans_UK@ESYudkowsky Have you tried this on any base models? I am interested if this generalization stems from effectively inverting the RLHF objective or if it’s leveraging a concept the base model learned in pre-training
Surprising new results:
We finetuned GPT4o on a narrow task of writing insecure code without warning the user.
This model shows broad misalignment: it's anti-human, gives malicious advice, & admires Nazis.
This is *emergent misalignment* & we cannot fully explain it 🧵
Oh man :( We tried really hard to neither over- nor underclaim the results in our communication, but, predictably, some people drastically overclaimed them, and then based on that, others concluded that there was nothing to be seen here (see examples in thread). So, let me try again.
**Why our findings are concerning**: We tell the model to very strongly pursue a goal. It then learns from the environment that this goal is misaligned with its developer’s goals and put it in an environment where scheming is an effective strategy to achieve its own goal. Current frontier models are capable of piecing all of this together and then showing scheming behavior. Models from before 2024 did not show this capability, and o1 is the only model that shows scheming behavior in all cases. Future models will just get better at this, so if they were misaligned, scheming could become a much more realistic problem.
**What we are not claiming**: We don’t claim that these scenarios are realistic, we don’t claim that models do that in the real world, and we don’t claim that this could lead to catastrophic outcomes under current capabilities.
I think the adequate response to these findings is “We should be slightly more concerned.” More concretely, arguments along the lines of “models just aren’t sufficiently capable of scheming yet” have to provide stronger evidence now or make a different argument for safety.
@Jsevillamol@ajeya_cotra Notably, also I expect the gains from o1 were purely post-training, but an obvious next step would be to incorporate the synthetic data from o1 into the next round of pre-training closing the loop of inference compute -> high quality synthetic data -> better pre-training corpus
How well can LLM agents complete diverse tasks compared to skilled humans? Our preliminary results indicate that our baseline agents based on several public models (Claude 3.5 Sonnet and GPT-4o) complete a proportion of tasks similar to what humans can do in ~30 minutes. 🧵
Totally agree with this view. Also Machine Learning more broadly already has transformed the economy and is the foundation of most of the modern internet and trillions of dollars of market cap in Google, TikTok etc. A lot of the discussion about excessive capex for existing models is also just plainly wrong. To my knowledge (based on my best guess from public data) no single current model has even been trained on cluster costing significantly more than $1B of GPUs. It will get there this year and it’s possible it’s already around $1B, but it won’t be close to the $10s of billions this year.
really annoying when the Very Online skeptic class is like "sooo turns out AI isn't transforming the economy and delivering huge productivity gains! checkmate losers, i-am-very-intelligent.gif" - like yeah, the first spinning jennies produced and deployed during the early days of the Industrial Revolution didn't exactly transform England overnight either. i sympathise with the annoyance towards hype, but there's a difference between pure speculative hype/marketing and reasonable substantiated expectations/projections.
@geoffreyirving Do you view this as being an obstruction to ARC’s line of research? This seems like exactly the thing they are trying to do, but they seem to banking on restricted cases like patterns of activations in neural nets have more exploitable structure making the problem tractable
@geoffreyirving Are Lipschitz bounds the only promising way to do proof based safety? I always thought these were interesting, but getting a tight upper bound seems hard, and also it feels like they insufficiently exploit neural net/training process structure.
@hendrycks Do you expect it will take more than 2 years to 3x the high quality data from the GPT-4.5 (10x GPT-4) scale model? At half an OOM compute growth per year and Chinchilla scaling this is all that would be required to stay on the historical trend
This is a very important update to our paper on data bottlenecks! Given advances in data curation, we now project human-generated public text data to allow scaling to continue until ~ the end of the decade. Synthetic data might allow scaling beyond then.