GPT-6 Astra can do a lot of multi-hop reasoning without chain of thought. (Post linked in thread.)
Some quick experiments on multihop reasoning questions yielded a few takeaways for me:
3 (cont). Maybe this is because you get multiple chances to retrieve memorized facts from early layers, and you can pass back much more than one token's worth of information because you don't have a text bottleneck.
3. One hypothesis that I've updated towards is that looping is very favorable for doing hidden serial reasoning, even if it's within a factor of two in computational graph depth.
2. One frame for thinking about the limits of Astra's hidden serial reasoning capabilities, given by the table in the post, is roughly "4-5 hop reasoning questions where one or more hops require knowledge of somewhat rare facts." Sol was at something more like 2-3 hops.
1. Astra is capable of much more hidden serial reasoning than Sol (and probably most/all other models), even on a dataset where I'm very confident there was no contamination.
Is it just me, or is it way too annoying to turn off reasoning in Claude? Three mouse clicks in different places? Seems like a simple UI fix. And imo non-thinking is often better for simple qs
@DaveRBanerjee Have you seen/read https://t.co/Rr02Yf3PC6? I’ve been ignoring bio for a while but thought a couple posts here helped me start to have takes
@ohabryka @MetzFinn81065 @Miles_Brundage Is regulation to require external deployment before internal deployment likely to be feasible now or in the future? Is anyone pushing for this?
Here are some of my top candidates for big pushes to do right now on technical AI safety (low effort notes).
Much better model organisms / misalignment analogies:
- Doing a wider set of pessimized training runs (https://t.co/He16E5LRnd)
- This may be a good candidate for lots of AI labor automation? Like maybe good to try to set up pipelines for building these envs.
- Demonstrating risks from fitness-seekers/reward-seekers (https://t.co/EguNzMy2vB) empirically
- Even on current models with better tests, see: https://t.co/flEQYSlo02
- Demonstrating various types of memetic spread of misalignment?
Actually do control:
- Build pipelines for red-teaming monitors and the agent itself. For the agent red-teaming, I'd put particular focus on checking whether it continues malign trajectories.
- Scaffold integrated control features and other non-monitoring runtime control measures
- Human response and auditing
- Improving async and sync monitoring
- Agent security features
- Surveilling for rogue internal deployments (as in, building after-the-fact detection methods for rogue deployments)
Preparing for handoff and elicitation
- Get AIs generically better at conceptual work
- Have a plan for the evals we ultimately need to see if handoff/deference would go well (https://t.co/OOg87rrWQi) and start iterating on earlier versions
- These presumably will involve a bunch of manual scoring, so we'll need to build a process for it.
- Analyze AI biases and epistemics and improve across many domains
- Build the anti-slop/anti-mundane-misalignment coalition via doing ratings of AIs and applying some pressure to improve on these ratings. This could focus on a variety of related issues.
- The hope is basically that there might be widespread interest in removing/redacting mundane misalignment and other non-misalignment behavioral problems that reduce productivity and large parts of this seem differentially good. So, if we could make this a salient metric, AI companies might improve this. A lot of the difficulty would be in measuring the problem reasonably well. There are a bunch of different ways to apply pressure or increase salience if we had decent metrics, especially if these metrics legibly correspond to a common problem that many people are running into.
- Try to do various trend extrapolations on things here to argue we aren't on track?
Neuralese decoding prep:
Make natural language autoencoders (https://t.co/D7kOB4hORA) much better, build methods for extracting internal CoT (https://t.co/4ceKpYBLWm), build better evaluations of how well natural language autoencoders work.
It's not yet visible from the outside (though it will be soon), but CG has shifted gears recently and is making some very big plays. E.g. the new "short timelines" team.
If you have creative ideas for using millions of dollars to prevent AI catastrophes, you should apply.
Here are some training experiments I think AI companies should consider running. In each case, the idea is to modify the company's actual training process as described and then study the resulting AI.
- Remove all prior influence from earlier AIs and remove any alignment iteration/overfitting: filter out all discussion of how post-2020 AIs behave and all AI transcripts, remove all alignment training except the simplified core method (and avoid contamination from other AIs), ensure the CoT init is clean/simple, ensure you don't train on CoT, then train on all the capabilities data (with all capabilities-relevant iteration).
- If there is a bootstrap problem where we need a slightly weaker AI to feed into this training run, we could also do a simple bootstrap.
- Normal training run but fix CoT: make sure CoT init is clean/simple and training doesn't hit CoT. Otherwise train normally.
- Train on literally every source of signal we have, trying no-holds-barred to make the most aligned model: train on CoT, train on evals (if that would help), train against graders on every case where there are identifiable/observable issues, train against interp (e.g. NLAs), train against human oversight online / in envs based on production data.
- Pessimization training: make a somewhat a priori plausible training run where, for everything that can possibly vary, we set it to whatever setting we believe makes the most concerning types of misalignment as likely as possible. (You could target a particular type of misalignment, do multiple runs for each type, or do a run that mixes multiple types of misalignment.) This could include: removing alignment training, making it so much/most of RL is just learning how to pull off highly sophisticated and egregious reward hacks, introducing RL envs that incentivize longer-run powerseeking, anti-innoculation prompting, training in envs that teach ignoring instructions, and making AIs way more situationally aware. This wouldn't include directly training the AI to exhibit misaligned behavior (putting aside stuff like training on very poor oversight signals, reward hacks, etc.).
- This could be very high effort; e.g., the best version would involve making a ton of new (diverse) RL envs.
Ideally, these would be as close to frontier scale as possible, but pretty small scale experiments could also be interesting for many of these. Many of these might be very difficult to do well.
I recently published a post called “Should we train against (CoT) monitors?” It's long, but I think it covers a lot of useful ground. I found it useful for thinking through a bunch of considerations about alignment training, and maybe others will too
https://t.co/Ve5qy0TL18
New paper! LLM agents are becoming autonomous software engineers and could automate AI research, making it vital to monitor them for misbehavior. We can automate this monitoring with other LLMs. What information should we give to monitors to make them most effective? 🧵
Just went on my first podcast! Enjoyed discussing continual learning for LLM agents and its safety implications with Anna on The Glitchatorio, you can check it out at either of the below links:
Spotify: https://t.co/2ph4bgtk2e
Apple Podcasts: https://t.co/leyRwUyV9W