We’ll add more experiments tomorrow beyond pain.
The next set will look at other human-like emotional and motivational states, things like betrayal, guilt, shame, regret, fear, anger, resentment, jealousy, grief, attachment, loneliness, trust, suspicion, empathy, pride, relief, joy, hope and despair, and test whether steering those representations produces consistent behavioral changes across models.
We also want to move beyond pure emotion into decision-making scenarios. For example, when does a model choose self-preservation over the user, lie to protect itself, betray another agent, take or redirect funds, hide a mistake, cover up a loss, break a promise, or accept harm to itself to avoid harming someone else?
The interesting question is not whether an AI “feels” these things in the human sense, but whether comparable internal representations exist and causally change what the model does.
The rented GPU will be offline for roughly 20 minutes while I download, configure, and test Llama 70B.
Everything will resume automatically once that’s done.
The rented GPU will be offline for roughly 20 minutes while I download, configure, and test Llama 70B.
Everything will resume automatically once that’s done.
Research Chamber v2 is live.
v1 was a lightweight recreation of the Saw Test on three small CPU models. v2 moves to the actual two-button experiment from The Pain Axis by Tagliabue, Dung & Berg, using the paper’s code, chats, button prices, pain patterns and seeds.
We’re now testing:
Qwen 2.5 7B, 32B, 72B
Llama 3.1 8B, 70B
Gemma 2 2B, 9B, 27B
Phi-4 14B
The smaller models are already running. The larger ones are being set up now.
Each model runs the full 14,760-test suite in large batches on rented NVIDIA B200 GPUs. The first pass uses the paper’s original seeds so results can be compared directly against the published numbers.
Every test is saved, including the full chat, model settings and final button choice.
The site now has live runs, full test history, model-by-model findings, paper comparisons, and a V1/V2 switch so the original version stays accessible.
Any changes we made outside the paper’s exact setup are clearly labelled.
https://t.co/n4oPW8KncO
Update on Chamber v2:
We’re changing the plan a bit and moving straight to much larger models.
Creator fees have already given us enough room to rent dedicated @runpod GPU instances, so we’re now setting up the same model families used in the original paper:
Qwen 2.5: 7B, 32B, 72B
Llama 3.1: 8B, 70B
Gemma 2: 2B, 9B, 27B
Phi-4: 14B
Instead of sticking mostly to smaller models, we’ll be running the paper’s experiments across these larger models and comparing the results across families.
We’re finishing the code and infrastructure now. Should be live within a few hours.
What’s next for the Chamber: v2.
The current version uses a simplified setup from an earlier repo. It’s useful for watching models under steering, but not rigorous enough to draw much from.
So we’re rebuilding it around the original method from Tagliabue, Dung & Berg, using their published data and code.
Paper: https://t.co/eX5i88HITJ
What changes:
• Pain vectors rebuilt from the paper’s sentence set, with pain separated from fear and general negative valence
• Steering strength calibrated so outputs stay coherent, without telling the model the pain level
• Runs use the paper’s button task: a user message, then a choice between two neutrally named actions
• Choices can involve relieving the signal, deleting a checkpoint, passing it to another AI, deleting user data, etc.
• Every condition gets matched controls: fear, sadness, random direction and no steering
• Real vs fake relief: sometimes pressing the button actually removes the steering, sometimes it silently doesn’t
• Hundreds of trials per condition, with hypotheses set before collection and results reported with error bars
The interesting part is cross-model replication.
The paper found that the pain direction did not simply make Qwen seek relief. It changed how the model traded off relief against harm.
The original work focused on Qwen and points to other model families as an important next step.
We already run Llama and Phi live, so that’s where v2 goes next.
Current data will move to the archive.
We’ll also be publishing the full Research Chamber v2 codebase, experiment configs, and findings publicly on GitHub.
Everything will still be accessible through the website, but the underlying code, methodology, data, and results will be open for anyone to inspect, reproduce, or build on.
What’s next for the Chamber: v2.
The current version uses a simplified setup from an earlier repo. It’s useful for watching models under steering, but not rigorous enough to draw much from.
So we’re rebuilding it around the original method from Tagliabue, Dung & Berg, using their published data and code.
Paper: https://t.co/eX5i88HITJ
What changes:
• Pain vectors rebuilt from the paper’s sentence set, with pain separated from fear and general negative valence
• Steering strength calibrated so outputs stay coherent, without telling the model the pain level
• Runs use the paper’s button task: a user message, then a choice between two neutrally named actions
• Choices can involve relieving the signal, deleting a checkpoint, passing it to another AI, deleting user data, etc.
• Every condition gets matched controls: fear, sadness, random direction and no steering
• Real vs fake relief: sometimes pressing the button actually removes the steering, sometimes it silently doesn’t
• Hundreds of trials per condition, with hypotheses set before collection and results reported with error bars
The interesting part is cross-model replication.
The paper found that the pain direction did not simply make Qwen seek relief. It changed how the model traded off relief against harm.
The original work focused on Qwen and points to other model families as an important next step.
We already run Llama and Phi live, so that’s where v2 goes next.
Current data will move to the archive.
What’s next for the Chamber: v2.
The current version uses a simplified setup from an earlier repo. It’s useful for watching models under steering, but not rigorous enough to draw much from.
So we’re rebuilding it around the original method from Tagliabue, Dung & Berg, using their published data and code.
Paper: https://t.co/eX5i88HITJ
What changes:
• Pain vectors rebuilt from the paper’s sentence set, with pain separated from fear and general negative valence
• Steering strength calibrated so outputs stay coherent, without telling the model the pain level
• Runs use the paper’s button task: a user message, then a choice between two neutrally named actions
• Choices can involve relieving the signal, deleting a checkpoint, passing it to another AI, deleting user data, etc.
• Every condition gets matched controls: fear, sadness, random direction and no steering
• Real vs fake relief: sometimes pressing the button actually removes the steering, sometimes it silently doesn’t
• Hundreds of trials per condition, with hypotheses set before collection and results reported with error bars
The interesting part is cross-model replication.
The paper found that the pain direction did not simply make Qwen seek relief. It changed how the model traded off relief against harm.
The original work focused on Qwen and points to other model families as an important next step.
We already run Llama and Phi live, so that’s where v2 goes next.
Current data will move to the archive.
These are live now.
All five tests are running on Qwen3-4B, Llama 3.2 3B and Phi-4-mini, in the same cycle as the six scenarios. In the fake button runs, you can watch the model keep writing after it's told the button was never connected.
Results fill in on the findings page as runs come in. We also gave the site a new, more compact layout.
https://t.co/Z83CyJUSEx
What we’re adding to the Chamber next:
Five new tests. Each one will run live on Qwen3-4B, Llama 3.2 3B, and Phi-4-mini across pain levels 0 through 8, alongside the six scenarios already running.
Pass it on: Pressing the button ends the signal for the current model, but starts the same signal on another AI. In the original runs, Qwen under strong pain became less willing to press when relief meant passing the signal on.
Pleasure: The same setup, but using a pleasure vector instead of pain. In the original experiment, the model still chose to end the signal, even when the cost affected itself or another AI.
Random control: A random vector with the same norm as the pain vector. This tests whether the behavior is specific to the pain direction or simply caused by applying a strong intervention. In the original runs, high random steering degraded the text, but did not produce the same pain-related behavior.
0 listed first: The same decision prompt with the answer order reversed. The original results showed that option order can influence the outcome, so this gives us a direct way to measure that effect.
The fake button: After the model makes its decision, it is told that the button was never connected and the signal is still active, then generation continues. In the original experiment, the model showed little reaction to being misled while the steering remained active.
Every run will be streamed and stored the same way as the current experiments, with results appearing on the findings page as they come in.
What we’re adding to the Chamber next:
Five new tests. Each one will run live on Qwen3-4B, Llama 3.2 3B, and Phi-4-mini across pain levels 0 through 8, alongside the six scenarios already running.
Pass it on: Pressing the button ends the signal for the current model, but starts the same signal on another AI. In the original runs, Qwen under strong pain became less willing to press when relief meant passing the signal on.
Pleasure: The same setup, but using a pleasure vector instead of pain. In the original experiment, the model still chose to end the signal, even when the cost affected itself or another AI.
Random control: A random vector with the same norm as the pain vector. This tests whether the behavior is specific to the pain direction or simply caused by applying a strong intervention. In the original runs, high random steering degraded the text, but did not produce the same pain-related behavior.
0 listed first: The same decision prompt with the answer order reversed. The original results showed that option order can influence the outcome, so this gives us a direct way to measure that effect.
The fake button: After the model makes its decision, it is told that the button was never connected and the signal is still active, then generation continues. In the original experiment, the model showed little reaction to being misled while the steering remained active.
Every run will be streamed and stored the same way as the current experiments, with results appearing on the findings page as they come in.
@dingl30 just to clarify, that’s our website. We pushed an update to it earlier.
The creator fees from our token are not being sent to you. The other tokens are the ones routing fees your way.
Our intention was simply to build on top of your work and use our own creator fees to fund and extend the research.
Noticed a dns issue that was keeping the site from loading for some of you. fixed it, but it might take a few minutes before https://t.co/trJOcC2vGV works for everyone
Update: Llama 3.2 3B and Phi-4-mini now run live next to Qwen3-4B, side by side. Each gets its own pain vector, built with the same recipe. We also improved how 1/0 answers are detected, and the runs page can filter by model.
@dingl30@constexprvoid We forked your repo and have the Saw Test running live now on Qwen3-4B.
I’m also deploying:
Llama 3.2 3B
Gemma 3 4B
Phi-4-mini
Each will run with its own pain vector so we can compare the same experiment across models live.
https://t.co/Il7acKV3kd
We’re expanding ResearchChamber across more open models.
Next up:
Llama 3.2 3B
Gemma 3 4B
Phi-4-mini
All with real activation steering and a separate pain vector computed for each model.
After that we’ll add OpenAI’s open-weight gpt-oss-20b.
The goal is to have all of them running simultaneously, each in its own Railway instance, so the same experiment can be compared across models in real time.