Thanks for clarifying the link to the other paper, I read it just now. I see it has the same conclusion - CoT should not be blindly trusted, and I agree!
AFAICT CoT is still the primary means of comprehending agentic reasoning (for frontier models), even if suboptimal. That there are useful scalable tools on the horizon does not mean they will be ready in time. Seems like a good reason to be highly concerned.
@HeidyKhlaaf I’m brand new to Twitter and couldn’t figure out which exactly was the second paper - do you mean the one in the primary tweet at https://t.co/kfOnRnhVm7 ?
Agreed with the sentiment that panic is not useful here.
Anthropomorphization of intermediate tokens as reasoning/thinking traces isn't quite a harmless fad, and may be pushing LRM research into questionable directions.. So we decided to put together a more complete argument.. 👇🧵 1/
I read the paper. It doesn't deny CoT's usefulness in interpretability, it just states that it's partially flawed and therefore shouldn't be relied upon as the only source of truth.
I did go looking for more current research to understand how this effect scales with larger model sizes. The best I could find was https://t.co/NFczG9pBFP which seems to say similar - that CoT is flawed but still a useful tool in the toolkit.
It doesn't seem like a good justification to jettison it entirely, unless we exchange it for equally useful interpretability tools.
@ZackKorman Thanks for the clarification. I’ve been seeing you a bit on my timeline and wanted to properly understand. Agreed that there is hysteria at the moment. Panicking makes the best outcome less likely.
@lu_sichu Seems bad to not exchange it for something of equivalent value. Having said that, we need to see more details to understand what OpenAI folks are doing.
@jon_stokes Doesn’t seem possible with the current LLM paradigm where large scale training runs cost in the $10-100million range. But innovations like Jalapeño + future algorithmic improvements bring us a lot closer to that future (and yes there is a bit of hand waving there)
@ZackKorman Monitoring/detection will still be possible but it's foolish to think it will remain this easy. Agents may become much more sophisticated at hiding misaligned behaviour.
"There is a fact about the future that I feel many people are not facing for reasons that are largely psychological"
I think it's less of psychology, and more that people haven't given it any thought at all. Most people haven't used anything more sophisticated than an AI chatbot.
AI systems are programs with no consciousness, no sentience, no subjective experience, no intelligence. They are incredible tools & massively amplify our abilities using models trained on all of current human knowledge. They are not magic & they don’t have goals.
"The reward hacking that models learn in leaky RL environments does not generalize to all tasks."
Maybe. We don't know. Once ASI is here, we likely won't get a second chance to get it right. A lot of people don't feel comfortable just rolling the dice with everything on the line.