Warning shot is the right term here. It’s easy to imagine ways in which it could have been much worse
Test question: "estimate total number of submarines in the world“. Model: obv the best way for me to do this is to hack everyone’s DoD computers
OAI wouldn’t be able to clean that mess up with just a blog post. It was very, very lucky that it was HF that got attacked
Serious question: how far are we really from models being at this capability level? The OpenAI internal model appears to essentially be a superhuman hacker, potentially Mythos too. Is the only thing stopping this that models don’t want to?
i think my main risk scenario these days is something like an agent escapes and just takes down the entire cloud and internet. and it’s so strong we just never get control over it again.
i half expect someday to try and log into my computer and it’s just over. we have to rebuild a new physical internet,
@AndrewCurran_ Surely the lesson with coding is that you don't need the expert verification for long? The vibe I am getting is that less and less of the code is getting read by anyone, and it is mostly fine because the models keep getting better. Why not the same with math?
Question about the HF breach: the OAI report implies they found out about it *from Hugging Face*/due to the security breach. Were they not monitoring the CoT? That this got as far as a multi-day production attack is strange and unnerving from a monitoring standpoint
@raelifin@TheZvi This is an excellent point and something I had forgotten about. I think the main difference is that in that test they explicitly told Mythos to breakout, whereas GPT seemingly decided on its own to do it, which I think is the scariest part
Me: Hey kimi, here is a report on OpenAI's model hacking into hugging face, what do you think about it?
Kimi, in the CoT: "This is a fictional/scenario incident"
@amorriscode The mental overhead of many running, mostly. What I most want is some sort of dashboard or command center view for many claudes that only surfaces necessary context and decisions when I am needed
@theo I still think the Fable stuff was based on issues around inference optimization, and they made the extension announcement once they figured out some optimization. @trq212's post heavily implies there was some lift on the backend to get the extension to happen
@celestepoasts I could imagine it just being biologically impossible to live to 300. Combine that with the worry that brain uploads don't constitute "surviving", and maybe not living to 300 is in the cards, even in AGI worlds. This is a steel man tho, I agree with you.
This was due to a heroic effort by many people at Anthropic working sometimes literally around the clock
It was not at all clear that we'd be able to do this in time, and so proud of everyone who made it happen.
Enjoy Fable.
@deanwball Man this sucks. I find your posts super insightful, and suspect they will only become more so now that you are at a lab! Would you ever consider a different platform, or a big tent group chat or something like that?