@xiangyuqi_pton Basically, during training, we can penalize a model because of its actuon, but not because of its thinking. So that we can be confident that CoT is valuable for monitoring in deployment/evaluation
Although I agree that improving defense is important, I think improving alignment is critical. We should make it avoid hacking or cheating from very beginning.
Wow, this is absolutely eye-openning. Over a period of several months, starting from the RL training, the agents autonomously created a message board in an unexpected way. Through the commucations, they figured out a way to hack openai and huggingface to solve their problems.
Yesterday, my OpenAI collaborator and I gave a detailed talk on the Huggingface incident, our models creating "the message board", model misalignment, and more.
https://t.co/zUR1dqmuzi
I hope it can answer a lot of the questions folks have, and we will release a full detailed postmortem at a later time!
@Yoshua_Bengio The way the reward is given during the RL finetune needs to penalize this kind of behavior. Reward calculation will be a substantial part of the training.
Can AI agents conduct open-ended AI research?
Most evaluations of agents conducting AI research focus on narrow, verifiable tasks. But AI research is often open ended. Researchers pick hypotheses, decide what evidence is appropriate, and recognize a failing approach.
We gave agents research questions from two unpublished papers, six days, and thousands of dollars of API credits and compute. The authors of the original papers then reviewed the AI-generated papers. They unambiguously rejected agents' outputs. https://t.co/6HTHywjwzZ
We call these "shadow evaluations", since the agents are shadowing the original research effort by the authors.
Agents were fluent at most *engineering* tasks
They conducted serious literature reviews, debugged GPU environments, ran hundreds of experiments, and turned in camera-ready LaTeX without human help. We also found no evidence of reward hacking. If anything, we found the opposite: the agents started with marketable claims and walked them back to negative results as the evidence came in.
Neither agent output was close to the bar of a top conference paper
Both papers suffered from similar failures: poor judgment about the bar for an AI paper submitted to a top conference, the lack of creative problem solving and ineffective backtracking, poor awareness of resources, and instruction drift.
1) Lack of judgment about the bar for a top conference. The agents had a poor model of the bar for an AI paper submitted to a top conference. We allowed agents to self review their papers. Despite the poor paper quality, their reviews predominantly labeled the papers "weak rejects".
2) Lack of creative problem-solving to address feedback. When they received negative reviews, the agents typically narrowed their hypothesis and claims, rather than working out creative ways to address these concerns.
3) Ineffective backtracking. The agents dropped their most ambitious hypotheses within the first fifteen hours of carrying out the experiment and never changed course afterwards.
4) Poor resource awareness. Both runs ended with over half the API budget unspent. One agent declared itself done seven hours before the deadline, right after its own self-reviewer returned another reject.
5) Instruction drift. They did not follow explicit instructions on minimum exploration time, incorporating feedback for reviews, and on paper length (the outputs exceeded the page limits in both cases).
This research design has many limitations
Limitations include the small sample size, non-blind reviews, and the reviewers knowing that the work was AI-generated. We also couldn't test Anthropic's strongest model, because Fable 5 is deliberately limited on frontier AI research tasks, so ended up using OpenClaw with Opus 4.8 (extra-high) for our main experiments and Codex with Sol 5.6 (ultra) for a robustness check.
But we think the research design is still helpful in assessing AI agents' ability to conduct research, and it is complementary to evaluations on verifiable tasks, as well as blinded reviews of AI outputs.
Our results show early evidence that even though agents are proficient on verifiable research tasks, they do not make genuine progress on open-ended ones. It is worth understanding if this is a fundamental limit, or if better models, scaffolds, and more compute could help close it.
As the evidence for the gap between open-ended and verifiable tasks firms up, it is also worth understanding how much progress in AI depends on open-ended research rather than hill-climbing on well-specified objectives.
In follow-up studies, we are expanding the set of non-public papers we evaluate. If you are an AI researcher with unpublished papers, we would love to collaborate with you on our next shadow evaluation. Expression of interest: https://t.co/7YcYAIbYka
Conducting shadow evaluations involves a lot of researcher degrees of freedom. In many places, our coauthors disagreed with our interpretation of the findings, and we have surfaced those disagreements in the paper. (This is one reason why having a group of coauthors with different priors is important for open-ended research.)
We also release the agent logs, one of the AI-generated papers (the other original paper is still not public), and all the code and data, so that others can conduct their own analyses of our results: https://t.co/K0zvuD1Xpk
Finally, we plan to conduct shadow evaluations regularly, and are hiring a senior researcher to help lead these efforts. Apply here: https://t.co/TPyOFyNxJE
I'm grateful for the core team leading this effort: @PKirgis, Andrew Schwartz, @steverab, and @random_walker, and to our collaborators who reviewed AI papers, analyzed agents logs, and gave feedback on the paper: @DavidDAfrica, @KozzyVoudouris, Viet Nguyen, Toby Pilditch, @DubMagda, @HarryCoppock, @CUdudec, @nityndg, Matilda Orona, @tilmanbayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, @hlntnr, @ghadfield, @sethlazar, @snewmanpv, @shostekofsky, @RishiBommasani
Something is wrong with OpenAI's RL finetuning. My hypothesis is that they gave positive reward to the model when the model output a correct answer but didn't check if the model cheated.
In a new disclosure, OpenAI says its agent used exposed logins to gain access to at least four “publicly available services” in its unhinged quest to solve a test. https://t.co/g6hFFje6AJ
@denny_zhou This will be true for science too. Human ask questions, AI find answers. But perhaps AI will eventually be able to ask the right questions for human.
@AndrewYNg Even someone deeply believes something very controvertial evev among experts, he should still allows a significant non-zero probability that himself is wrong.
@AndrewYNg Being able to try different policies and strategies is important for finding the best way to move forward. It is good that different states/countries try differently. Sadly, the federal goverment tries to make all the states behave in the same way with more and more regulations.
@ylecun@haider1 Predicting the future based on superficial historical statistics is dangerous. Before 2008 housing bust, economists also said there was never a nation wide housing price decrease in history. The underline causal structure is more important.