Geoffrey Hinton and Andrew Yao led an international convening of researchers for the fourth International Dialogue on AI Safety (IDAIS) in Shanghai, China. They shared and discussed recent studies and evidence that some AI systems today already demonstrate the capability and propensity to undermine their creators’ safety and control efforts. 🧵
1. Mandatory safety assurance from frontier developers – rigorous safety and security evaluations plus pre- and post-deployment monitoring for high-capability models.
2. Global verifiable red lines – focused on the behaviour of AI systems and committed to through an international coordination body.
3. Safe-by-design R&D – moving from reactive approaches to safety issues in current systems to proactively building systems that are safe by design.
The attendees drew upon a range of evidence suggesting that AI systems can detect when they are being evaluated, pretend to be aligned, and resist being shut down. It warns that today’s frontier models already show both the capability and propensity to undermine the safety and control efforts of human operators. In response, it urges:
1/ Consensus statements from the International Dialogues on AI Safety (IDAIS) are more than just words on paper – they are a roadmap for action. 🧵 on a guide to action 🎬
We don't have to agree on the probability of catastrophic AI events to agree that we should have some global protocols in place in the event of international AI incidents that require coordinated responses. More here: https://t.co/FrKyLXJLYf
In her piece for @nytimes covering our recent International Dialogues on AI Safety event, @megatobin1 notes our events "are a rare venue for engagement between Chinese and Western scientists at a time when the United States and China are locked in a tense competition." Link below
@Yoshua_Bengio@yaqinzhang@berggruenInst 13/ The IDAIS program is funded and supported by the Safe AI Forum, a US non-profit organization dedicated to fostering international cooperation on AI safety. Read more about us at https://t.co/e9m92LlYRp.
Leading computer scientists from around the world, including @Yoshua_Bengio, Andrew Yao, @yaqinzhang and Stuart Russell met last week and released their most urgent and ambitious call to action on AI Safety from this group yet.🧵
@Yoshua_Bengio@yaqinzhang 12/ This event in Venice was also graciously co-hosted by @berggruenInst, which you can read more about at https://t.co/McJgR93QAR