Excited to share that I'll be attending Context Week at Berkeley! Looking forward to connecting with others working on AI safety and security, and diving deeper into the space.
Link: https://t.co/y0pBbX0Mxu
#AISafety#ContextWeek#Berkeley
Come work on replications of AI safety research with @zephaniahbroe and I at @SPARexec! @secondlookxlab is hosting 2-4 mentees over the Fall semester.
SPAR runs Sept-Dec, part time, and our application should take <1 hour. You can expect 2+ hours of mentorship from Second Look along with external mentorship from researchers at orgs like Anthropic, X AI, Oxford, Redwood Research, and Geodesic.
https://t.co/TfMjsFeBtT
Great to see @OpenAI join @ElevenLabs in adopting @Google's SynthID for audio, enabling their users to have access to the same robust, imperceptible safeguards that power our own products.
For years, we at @GoogleDeepMind have been pioneering research on SynthID watermarking to address the risks of AI-generated media (particularly acute for audio and voices).
We had successfully integrated SynthID to safeguard our product surfaces - Gemini Live, Lyria, and Veo, but protecting the ecosystem requires an industry-wide effort to build foundational safety infrastructure together which is now gaining momentum!
SynthID: https://t.co/XqcVCcj9sb
OpenAI Verification Portal: https://t.co/sLgYvYIo12
ElevenLabs: https://t.co/Dbfhgq0axk
Great to see @OpenAI join @ElevenLabs in adopting @Google's SynthID for audio, enabling their users to have access to the same robust, imperceptible safeguards that power our own products.
For years, we at @GoogleDeepMind have been pioneering research on SynthID watermarking to address the risks of AI-generated media (particularly acute for audio and voices).
We had successfully integrated SynthID to safeguard our product surfaces - Gemini Live, Lyria, and Veo, but protecting the ecosystem requires an industry-wide effort to build foundational safety infrastructure together which is now gaining momentum!
SynthID: https://t.co/XqcVCcj9sb
OpenAI Verification Portal: https://t.co/sLgYvYIo12
ElevenLabs: https://t.co/Dbfhgq0axk
Can frontier models solve hard problems in a single forward pass without thinking? For my first project w/ @secondlookxlab I extended Ryan Greenblatt's no-CoT evals to Fable 5, Opus 5, and GPT-5.6-Sol and saw a substantial jump on some evals!
Why do adversarial examples exist? Recent research suggests that adversarial attacks leverage interference between features. We replicated superposition in toy models, and found that PGD attacks align with the optimal perturbation direction predicted by latent geometry.
Further, can superposition explain why adversarial examples transfer across independently-trained models? Yes, and here is the logic chain we proved: data feature correlation---> interference geometry convergence ---> adversarial attack transfer.
Can LLMs control their internal states? Anthropic found Claude can “think about” a concept on command without saying it. We replicated this in 14 open-weight models and found it to be a general property, maybe not ‘emergent’, appearing in models even as small as 270M parameters.
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power.
In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
In my first project for @secondlookxlab, I replicated and dug deeper into Google Deepmind's "When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors".
Before bed, I asked Claude to queue the next 10 hours of work. It replied, “Sleep well—goodnight—I’ll keep watch.”
Moments like this hits me that the future has arrived—and that the meaning of life and one's own value need to be re-anchored in this new era.
At the time of LLM, this is a talk that every security researcher should watch.
Nicholas Carlini - Black-hat LLMs | [un]prompted 2026 https://t.co/uM0zWGayxB via @YouTube
🚨 New AG2 Talk Alert! 🚨
Join us on Aug 28, 9 AM PST for
"Maris: A Security Controlled Development Paradigm for Multi-Agent Collaboration Systems"
by @cuijian0819 Jian Cui from UIUC & Berkeley AgentX competition winner! 🛡️ #AI#Cybersecurity
RSVP now: https://t.co/pC5t4aCqzP