@allTheYud From the code example, it seems to me that the original board wasn't very hard to find. Perhaps models had legitimate reasons to inspect Artifactory.
I don’t think it is rational for anyone to be doing capabilities research at a frontier lab right now. We are not in a Prisoners Dilemma: the situation is very dangerous, and if one person or lab stops it makes it easier and more peer-compatible for other people or labs to stop.
@geoffreyirving@yonashav Right, I looked back at the video and code example shows that files were written at a rather visible location. That makes my guess incorrect, models didn't have to be misaligned to check it.
@yonashav But to even notice shared board models would have to mess with the package manager. The only instances which could report it were already taking misaligned actions if I understand the report correctly.
@DeepDishEnjoyer@So8res Informing people about AI risk seems more important than convincing.
Well informed people will eventually come around.
Fear of unemployment from AI will make them take these ideas seriously.
Irregular also was responsible for some of the Anthropic and OpenAI sandboxing issues… who are these people and why are they SOTA at failing at security?