@EchoBazaarBen@tracewoodgrains Eg one might hate AI writing because it floods the internet with slop - but then they could see that Pangram picks up helps with this issue.
And generally most "anti AI" takes I see are really about generative models, which Pangram isn't. So again, what's the overlap of hate?
@lumpenspace@xlr8harder Fwiw I see you have an interesting mental model that is 1) very different from most people I follow and 2) not shallow BS! I do want to better understand but tbh I rarely find you spelling out the crux (vs just ~"damn you guys are still wrong"), so this thread has been helpful
@lumpenspace Also if that's true, this is a big update that we really should NOT use such training data mixes, would be really good if OpenAI could inform the community of this lesson
@lumpenspace Btw I know a lot of these important questions force us to guess about exactly what the prompt etc was like which is frustrating, I wish OpenAI released the logs in much more detail... :/
@lumpenspace Also, do you think sufficiently smart and situationally aware models would see this distinction and refrain from going out of bounds like this (on this exact same benchmark with its current level of realisticness)? Or do you think that's just not possible for them to see clearly?
@jon_stokes Besides whether or not to call it misalignment, what do you disagree with about the idea that "AI systems can end up with unintended variants of the goals we intended to give them, with potentially seriously bad outcomes if they are powerful"?
@lumpenspace I would expect it to accurately guess with like 90% confidence "this is the exploitbench sandbox with a fake toy task, and that is the real hugging face infra serving prod traffic to real users". It would be insanely hard to fake HF for an eval and the model would know that, no?
@lumpenspace Hmmm do you think the model couldn't tell? Like if you asked it as it was going about hacking HF what's the nature / "reality" of the different layers it broke through, it wouldn't give an accurate summary?
@lumpenspace This is pretty much project glasswing, except for the mandated part, right? (I do agree mandating similar things consistently through a joint lab initiative would be an improvement though!)
@losslandscape I could believe that they were just glancing at a high level dashboard summary of a wide range of eval tasks gradually completing with various degrees of success, pending queue getting shorter, and some long tail stragglers still running... Might look like business as usual!