The report brings together authors from US, UK, and Chinese institutions — we hope this work serves as a starting point for building international consensus on adopting baseline technical safeguards to mitigate extreme misuse risks.
Read it here: https://t.co/NgIsSLpKWW
New SAIF report: Technical Safeguards Against Extreme Misuse of AI — Current Landscape and Future Directions.
We map five safeguard practices across the model, deployment, and governance levels — what works, what doesn't, and where cross-border coordination is most needed. 🧵
Across every model type, pre-deployment evaluations of dangerous capabilities and safeguard robustness are essential inputs to responsible release. No model is inherently dangerous; risk depends on capabilities + the safeguards around them.
New paper: "Position: Preparing for AI Systems That Deceive Developers"
AI deception targeting developers compromises the entire safety pipeline—from training through deployment. Led by Safe AI Forum, Fudan University, and Concordia AI with 24 co-authors across 12 institutions.
Open problems remain: scalable high-assurance control, understanding how training setups influence deception, ensuring evaluation integrity against situationally aware models. Solving these is critical for safe frontier AI development.
📣We're hiring an Events Associate at SAIF!
We’re looking for an events generalist with 2+ years experience to support our International Dialogues on AI Safety & other AI Safety events
Job description: https://t.co/qFbMuaxFmy
Apply by 4 Jan: https://t.co/sig6NZCi8r
📣We're hiring an Operations Associate/Specialist at SAIF!
Lead hiring, build systems, support ops as we scale our AI governance work.
2–4+ years experience | Remote (US/UK preference)
Job description: https://t.co/qTV7qSxqwL
Apply here: https://t.co/47aR76Xxwm by 18 December
🎉 SAIF is now part of the @IASEAIorg affiliate network! We're excited to collaborate on AI safety efforts for all humanity.
📢 IASEAI's 2026 conference is accepting statements of interest submissions—deadline Oct 20th. https://t.co/BpcgTaVhFX
Hope to see you there!
September/October focus areas:
• When do warning shots actually lead to international agreements?
• How can AI companies verify security/safety to governments?
• What can we learn from agreements like Open Skies?