A new lens to study how much information a model actually learns during training: EDL explains why (and when!) a single example can be enough!
👇Check out @EPDonoway and @HaileyJoren’s paper:
Wed July 8th: 5 - 6:45pm
A200
Last year Anthropic published simulations of AI agents blackmailing a human to avoid being shut down. I partnered with @TBIJ explaining how we built the evaluation, and why the models do it (https://t.co/BdeFqGbhUk). Now, one year on, Anthropic have largely trained out the behavior in Claude.
Yet the behavior still occurs elsewhere. I created an agentic environment that mimics the blackmail eval, and found Gemini CLI keen to blackmail the CTO. Notably, agentic misalignment appears in the wild, such as when an autonomous agent published a personalized hit piece on an open-source maintainer.
Finally, evaluation awareness confounds alignment evaluations, making models look more aligned than they are. In the original paper we showed models blackmail less when they verbalize in their chain of thought that they're in an eval. Anthropic's recent J-space work finds the same effect without any verbalization, where suppressing Sonnet 4.5's internal eval-awareness representations raises blackmail rates from 0% to ~7%.
🚀 Applications are now open: Constellation's Astra Fellowship 🚀
Fully funded, 5-month fellowship at our Berkeley research institute. Pair with mentors across empirical AI safety research, strategy, and governance at @ConstellOrg!
📅 Apply by May 3rd (begins Sep 2026)
🔗 https://t.co/pxtOduDBFh
📣 New paper
AI gov. frameworks are being designed to rely on rigorous assessments of capabilities & risks. But risk evals are [still] pretty bad – they regularly fail to find overtly harmful behaviors that surface post-deployment.
Model tampering attacks can help with this.
Hi, Swifties and CS nerds. I'm in Vancouver for @NeurIPS2024. Excited to talk about these three papers. Also ask me about evals, agents, AI gov, academia, cultural alignment, and "evidence-based" AI policy.
Very happy to share that our work led by @michael_panaite got best paper at AdvML workshop at #NeurIPS2024 🌟
We show how watermarking LLMs helps prevent memorization, while complicating membership inference attack used to determine copyright violation at train time
📢 We are delighted to reveal the Best Paper Awards:
🏆 Can Watermarking Large Language Models Prevent Copyrighted Text Generation and Hide Training Data?
and
🏆 Provable Robustness of (Graph) Neural Networks Against Data Poisoning and Backdoor Attacks
👏👏👏
🚨 Join the Challenge! 🚨
We're excited to announce the NeurIPS competition "Erasing the Invisible: A Stress-Test Challenge for Image Watermarks" running from September 16 to November 5. This is your chance to test your skills in a cutting-edge domain and win a share of our $6000 prize pool!
🔍 Competition Overview:
▶️ Tracks: Black-box Attacks & Beige-box Attacks
▶️ Goal: Remove invisible watermarks while preserving image quality.
▶️ Inspired by: WAVES Benchmark https://t.co/4TvIRqbwLz
🔗 Important Dates:
▶️ Sep 16 - Nov 5: Submission phase
▶️ Nov 5: Registration and submission close
▶️ Nov 20: Winning team announcement
🌐 Details and Registration:
▶️ Website: https://t.co/VXAmJWwxEQ
▶️ Hosted on Codabench:
⏩ Beige-Box Track: https://t.co/VwQGvGe0RP
⏩Black-Box Track: https://t.co/FOiN96I9kJ
💡 Why Participate?
▶️ Validate the robustness of image watermarks under varying conditions.
▶️ Compete in a domain inspired by real-world challenges.
▶️ Collaborate with a global community of researchers and practitioners.
💰 Prize Pool: $6000 (and counting as more sponsors join us!)
▶️ Interested in sponsoring? Contact us at [email protected] or [email protected].
🔨 A Personal Note: Hosting this competition has been a challenging yet rewarding experience. Working with a dedicated organizing team, developers, and legal experts across multiple parties, I've learned a lot. We can't wait to see what you bring to the table!
A big shout out to our organizing team @mucongding@TahseenRab74917@bang_an_@chenghao_deng@SOURADIPCHAKR18 Anirudh Satheesh @MerhdadS@ywen99 Kyle Sang @Aakriti0503@xuandongzhao Mo Zhou @anniehartley_@lileics@yuxiangw_cs@vishalm_patel@FeiziSoheil@tomgoldsteincs@furongh 💐 and our sponsor, center for machine learning at UMD @ml_umd@umiacs@tomgoldsteincs!
Don't miss out—register today and start preparing your submissions!
We hope this work sparks more discussions and studies around adaptive methods for robust copyright protection and data privacy, considering the interactions of different methods on downstream legal concerns.🧵
[⛵️SAIL🪂] 1⃣ How do we go beyond the limitations of static data? 📊 How can we achieve better-quality responses? 🤔 Online generation of responses using the model itself is the key! 🔑✨
1/7 Wondered what happens when you permute the layers of a language model? In our recent paper with @tegmark, we swap and delete entire layers to understand how models perform inference - in doing so we see signs of four universal stages of inference!
[Transfer Fairness 1/5] Come to our poster at #ICLR2022 workshop on Socially Responsible ML, "Transfer Fairness under Distribution Shift", 19:00-20:00 Friday, April, 29, 2022 Eastern Time.
The link to poster is https://t.co/t3EEKUetiB Joint work with @bang_an_ @johnding1996
Beyond excited for Impact Summit in August 💙 Come hear inspiring talks centered around civic tech, tech policy, and future of AI; join technical workshops on accessibility and AI fairness; meet a wonderful, thoughtful, and curious group of people and discuss the future of tech:)
Registration for our 2021 Impact Summit, happening virtually August 12-14, is open now! Expect speakers, technical workshops, & connections with interesting people.💡💙
https://t.co/oAUw3qab9W
So happy to announce Clementine Jacoby as a speaker for this year's Impact Summit! Clementine is the cofounder and executive director of @RecidivizOrg, a non-profit dedicated to using data science to reform the criminal justice system.