New paper https://t.co/WoMwRuVkos: we fine-tuned GPT-4.1 to be "evil", then realigned it. Evil models rate themselves as harmful. Realigned models do the opposite - without seeing any examples of their behavior. This shows LLMs can introspect and possess a form of self-awareness.
New paper (https://t.co/LCxYmNRXSw): models often verbalize evaluation awareness (vEA), and safety researchers worry this means models are gaming the eval. We injected/removed vEA in the CoT. It barely canged behavior, showing that vEA may be a weaker safety signal than assumed.
New paper in @NatureComms: while AI fairness research usually focuses on humans, we show that LLMs also discriminate against animals - a bias that has barely been studied so far. Thanks to my co-authors Monika Jotautaitė, @brewster_tw, and @LuciusCaviola.
https://t.co/hf6n2UFfXs