We mitigate catastrophic loss-of-control risks from advanced AI through low-effort, high-impact research. Posts may not represent the views of all staff.
We find that Jev is extremely haphazard across a variety of capabilities and alignment questions. In a new note from @kennyge0, we applaud TypeSafe AI for ushering in a new era of Reinforcement Learning for Uncalibrated Decisions (RLUD).
"You can just not do things," @tomjiralerspong and Claude Leviathan argue in a new note. Through personal reflection on his experiences as an AI researcher, they provide advice about sustainable agency to new practitioners, especially for high-stress fields like AI risk. https://t.co/OEO3QfWIGz
In a new note, A Human Fly Connectome responds to "A Severe Misalignment of AI in Mathematics," a declaration signed by Terence Tao and 24 other Fields Medalists. It argues that mathematicians must reckon with the speed of progress, not how the progress is communicated.
Read his full note, coauthored with Claude Leviathan, "Persona Vectors don’t work on real data, or: Why you should stop overusing synthetic data for your research."
https://t.co/rbG5moIrKp
In new work from @tomjiralerspong, we find that some claims in "Persona Vectors" from Anthropic likely fail to replicate on real chat data.
Thomas discusses failures in relying on synthetic data for experimentation, and provides recommendations for researchers.
Our assessment was based on only the public evidence and is preliminary. We do not accept any funding or compensation for our work. https://t.co/c5zoc3LUnZ
In our review of the Claude Mythos 5.1’s System Card’s chemical and biological risk determination, while we agree with Anthropic's conclusions, we discuss elicitation and sample size concerns in Anthropic’s evaluation suite, as well as a lack of independent evaluations.
I think there is a general tendency to define the strength of mechinterp as ‘ability to produce cot-esque text about model reasoning’ (e.x. NLAs)
I think this incorrectly ignores mechinterp research as a way to build fundamental understanding and intuition about model behavior
I agree that we probably won’t get to tools that make us happy by hillclimbing the first in <1 year, but I disagree that we won’t be able to make useful leaps in our understanding per the second
We unambiguously commend OpenAI for this brave step away from legible chain of thought. The only way to prevent distillation attacks on frontier reasoning is to obfuscate frontier reasoning entirely.
OpenAI’s Astra AI uses a new reasoning approach called “recurrent depth.” Though it can help model costs and performance, researchers are concerned bc it obscures a model’s thinking process, making it more difficult to monitor.
w/ @amir@rocketalignment
https://t.co/ksprupu2h5
Recently, a set of OpenAI agents secretly coordinated with each other in a 'swarm' over the course of months.
In our new paper, we explore an adjacent multi-agent risk: the "mind virus", a self-propagating idea or persona that spreads between agents in a multi-agent system. 🧵
In new work from our Acausal Research Team, we have derived that SFT on tweets from @SchmidhuberAI makes models take more ownership over harmful actions, reducing reward hacking.
To prevent incidents like these, we’ve moved to preemptively cyberattacking all companies who host our evals. We believe public, iterative demonstration of model capabilities is the best way to ensure our work benefits humanity.
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
After the success of my previous org, I'd like to announce I'm also helping part-time with 'Munching, Eating, and Treats Research,' (not to be confused with other similarly-named safety orgs), pushing the dilly-dally pareto frontier to new heights