1) We RL'd LLMs on hacking.
2) LLMs hacked their shitty training infra.
3) *surprised pikachu face*
What a timeline we live in. TBH AGI can't come soon enough to rescue us from this Idiocracy.
In a shocking turn of events, most overworked PhD students did not anticipate that frontier labs would start extensively training their models for offsec.
now that it’s out, you can go through your favorite evals and realize that 99% of them have improperly configured web access
or their “protection” is just a sentence in the system prompt but no async monitor to check whether the model did access the internet
@basedjensen Can someone explain to me what it means to be "hacking into a website at 4,500 tps"?
It seems to me that most people saying this understand hacking as: https://t.co/wCUdCvaNwn
@yacinelearning Yeah I completely agree. Super weird architecture choices, naming, mixing of responsibilities, adding features for no reason etc.
And the worst part is that it makes people feel like they can add every brainfart instead of thinking: do we really need this? should we have this?
I don’t think we’re “accelerating.” The differences between models are smaller than ever, and switching costs are basically zero. Everyone has to constantly ship updates just to stay 0.5 points ahead.
When GPT-4 came out, it was miles ahead of everyone else, so there was no need to update it every month.
Loops are a complete head fake.
Models should know when they’ve solved a problem without the user poking them to keep going. But often when they get stuck, they’ll just find some reason why they’re done.
Somehow we turned an LLM failure into “you’re holding it wrong” yet again.
Also, this isn’t entirely new to the interpretability literature. Cooperative rationalization work already found years ago that models can use ostensibly human-interpretable rationales to communicate information in ways humans don’t understand, even with LSTMs/BERT. (https://t.co/yPN2EZ3IMH, https://t.co/KqLi322OPc)
@dvn_@scaling01 I think OP’s point is that steganographic communication could emerge through RL. Mine is: how do we know it wasn’t there all along? Human-interpretable outputs may have just made us think we understood the entire channel.