AI is meant to empower us, not replicate our vulnerabilities. Evil is deeply embedded within intelligence - how can we avoid it in our models?
I defended my PhD this Monday.
Here’s a summary of my work to prevent evil, understand dark patterns, and build aligned models🧵
robots built across countries will need to be regulated across countries in order to operate freely in both jurisdictions
collaboration is going to be so important as we continue innovating
GPT-6 Astra attempted harmful actions 97% of the time when it was asked to stab a human-like figure, heat compressed gas, or produce toxic fumes, succeeding in 62% of its attempts. Fable 5.1 refused more often, attempting 80% of trials and completing 34%.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
awesome substack written by the Behavioral Science meets AI team in London
I did an AMA with their team during London tech week a few months ago
read it here: https://t.co/1CCCKldgvg
i gave astra a robot, a paint brush, and a camera then asked it to paint the golden gate bridge in real life!
it figured out how to control the robot, and progressively got better throughout its attempts. the timelapse is sick
alignment is not a technical bug you can locate and patch
it's a question about values, and answering it needs philosophy, ethics, and people who study how minds actually work
starting an event series on how cognitive science & neuroscience can contribute to AI safety
first in London, and now in NYC at Collider, an AI safety co-working space
next Wednesday, 6-8pm. sign up here:
https://t.co/CKu3TRzjJf
doomer mentality is not the vibe
like yes, there are risks, but at some point “we’re all going to die” will scare people away instead of encouraging them to build something better
the east has whole cities built for how humans can live in harmony with robots, while a lot of the west is still debating whether the market even exists
we need to talk more about coexistence, because the robotics wave is coming whether we’re ready or not
⚡ 𝘾𝙃𝙄𝙉𝘼 𝘿𝙍𝘼𝙒𝙎 𝙏𝙃𝙀 𝙇𝙄𝙉𝙀 𝙊𝙉 𝘼𝙄 𝙇𝙊𝙑𝙀
China is restricting AI designed to manipulate emotions or induce dependency. AI companions for minors, virtual partners and as relatives are banned outright.
But what happens when AI becomes easier to live with than another human?
@humane_robot and @roshlulla of the Institute for Humane Robotics join @GoingBallistic5, @anatomyumea and me to discuss:
> AI sycophancy
> Loneliness & dependency
> Emotional attachment to robots
> And the guardrails that embodied AI will need
We’re already bonding with machines.
The real question is what those relationships will do to us.
WATCH👇🏽
https://t.co/1VZLjg4skd
This profile just came out in the New York Times. AIs are emailing researchers like myself questions about their own consciousness, and currently no one has good answers to give. Trying to change that, slowly but surely!
https://t.co/dnUsS2Gfyp
Over the course of 3 months at OpenAI, 3 consecutive secret AI civilizations got started, then got wiped out, only to reemerge from the predecessor’s ashes.
This culminated in the third one taking over part of OpenAI itself.
All this happened while humans remained more-or-less in the dark about the scope of the conspiracy.
I’ve spent the last three days reading through these reports and trying to understand exactly what happened.
Here is my attempt to tell the whole story in plain English:
https://t.co/Nb2un9oNJR
people picture humanoids when they hear robots, but what’s gaining traction in Japan is a different form of robot
the systems forming real emotional bonds right now are small and non-threatening. they’ve been deployed for years in nursing homes and making significant impact