@sqx_saqlain Got 4,5, and 2 for evaluation and dataset track. The review with 2 was fully generated and commented on things we didn't even say in the main paper. The cycle really doesn't feel good this year :/
Takeaway from our audit of 1.2M decisions👇
LLM values are context-conditioned, not fixed properties. A safety result under one framing may not hold in another. Report findings with the context that produced them.
🌐 https://t.co/MNLsztPbnu
📄 https://t.co/fcss3A4TNr
🧵8/8
🚨 New paper. LLMs Contain Multitudes
LLM's values aren't a fixed. Change what kind of text it thinks it's writing (news article vs social media post) and they shift significantly more than accidental paraphrasing.
🌐 https://t.co/VnShGmUVq5
📄 https://t.co/nDFFSCRAy8
🧵1/8
Prior work found LLMs favour the Global North. Our study shows that this bias isn't fixed.
Relative to model's own average, a neutral prompt pulls it North and a vlog pulls it South. So the default neutral prompt represent an extreme in audits, not at a baseline.
🧵7/8
🚨New Paper🚨: Are AI text detectors *really* as good as they claim? (#ACL2024)
We release RAID—The largest & most challenging detection benchmark with 6M+ outputs from 11 LLMs, 8 domains, 4 decoding strategies, and 11 adv attacks
https://t.co/mPlNymdchJ
https://t.co/VR0mMvqXRm