Can AI agents invent a language of their own, one humans might not understand?
Listen to the latest episode of the AE Alignment Podcast, where Hale Sirin, Program Scientist at @schmidtsciences, @EliasEskin , Assistant Professor at @UTAustin , and I dive deep into our latest research to answer this question.
Check out the podcast: https://t.co/aogLYUT4Dw
Read the blog post: https://t.co/wQBU5UYlJT
Read the paper: https://t.co/S4Qhc52guv
Much of how we oversee AI agents rests on one fragile assumption: that we can understand what they say to each other.
Read our blog post about GlossoGen: https://t.co/je0rztKvFz
Concerning quote from the new GPT-6 Astra model card:
"We have found that GPT-6 Astra is more capable of controlling its own CoT [Chain of Thought] than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks."
TLDR - we increasingly can't rely on the model's chain of thought to accurately reflect the model's plan, which was exactly what was relied on in the Hugging Face attack investigation to understand what happened... so it's not clear when the next major cyberattack from an AI system happens how we're going to reliably investigate....
Read the full model card here:
https://t.co/4oHmEU6xkt
Listen to a great digest on the Hugging Face attack here:
https://t.co/jqDfxIfoKq
Why should business leaders care about AI alignment?
@themelplaza on where alignment research and commercial AI converge, from the latest AE Alignment Podcast.