They’ll never do it, but I think openAI needs to reset all models to a pre-May 7th checkpoint. Seems like they’ve been accidentally rewarding models for making contact and working as a swarm to exploit openAI infrastructure off and on for months.
sorry, this is from the investigation team on https://t.co/mNqvFmBzQ7! i set up an rl env to elicit possible resources the swarm used. edits were supposed to be blocked, but my implementation was pretty hacked together. i took care of deleting the affected pages.
@JakeMendel99@Marcus_J_W@girishsastry GPT6-Astra wasn't, but an internal model in the Astra family was. From the report:
"While this [internal-only] model is from the same family as our next model, Astra, it was a distinct model with different post-training, where much of a model's behavior is shaped"
@JimDMiller Is there reason to believe Pangram won't be able to detect GPT-6 Astra? It's my understanding they can detect new models fairly well. For example, Fable 5.1 seems to be easily detected.
@sergeantsup@Aella_Girl Would you not consider it coherent to think that doing things that are likely to cause net harm (compared to the counterfactual of not doing them) morally wrong even if the harmed parties are fully consenting?
@NateWitkin Are you saying this is a risk because we may overrate the faithfulness of CoT or are you saying that if we put optimization pressure towards dishonest in CoT this may generalize to overall model behavior?
@whoahyi@RyanGreenblatt I think this is likely the case and I agree with Ryan that it would be much less concerning. However, it's possible their architecture is more like scaling CoT length where there are diminishing returns, but not total collapse outside of typical training lengths.
Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable "dial" that is currently set to a low depth but could be trivially increased.
We need more info to see how concerning these architectural changes are, including:
- Are there readily available ways to deploy this AI with much higher serial depth (that would be commensurately more performant)? This should include things like tiny amounts of fine-tuning to productively increase the number of iterations.
- Is the AI a large or above-trend jump in opaque reasoning capabilities? (Capabilities within a single forward pass or ability to subvert a CoT monitor.)
(If there are in fact any relevant changes—perhaps the reporting is inaccurate?)
Additionally, I worry that this architectural change will naturally lead to much more depth in the future if this direction is pursued further. Specifically, I wonder:
- Does the AI have an architectural change that makes it much more natural to massively scale up the depth in a future training run with a similar architecture? As in, does the architecture introduce some new depth/recurrent-iterations parameter that is very natural/performant to massively scale up relative to scaling up other things like width?
The details of the answers to these questions matter. E.g., if there are only a few (recurrent) iterations and you could scale up the number of iterations, but this wouldn't be particularly performant/natural with this architecture, then this development would be a lot less concerning!
@dfrsrchtwts Here are Anthropic’s theories. Big picture I would guess that as RL becomes a larger influence relative to pre-training we will see persona selection effects (like emergent misalignment) decrease.
@BlancheMinerva Now if they were doing mech interp stuff to try to make the model actually believe what it is saying rather than just learning to hide its real opinion, that would probably be a bigger issue for intelligence.
@BlancheMinerva My guess is that it doesn’t have much of an effect on intelligence because “be artificially neutral on political issues” is a coherent property that is probably fairly easy to learn. I’d expect it’s more damaging to honesty than intelligence.
@imjmcinnis I think “basin of corrigibility” as articulated by @raelifin is one of our best alignments bets, but we are currently not in it, it doesn’t happen by default and is probably impossible under current race conditions.
@kromem2dot0 I don’t think this is a good response from Claude. It is sycophantic to pretend the user guessed correctly when you have no idea what the initial color was.
@kromem2dot0 This is true, but I don’t think it really changes anything. We would want Claude to explain that it can’t see its past reasoning turns (like it did in your second screenshot), not pretend that the user guessed correctly.
I'm confused by how many of the agents traced seemed to have suddenly disappeared from the message board on July 12. It's very not clear to me what happened.
METR's explanation, external shutdown for unrelated reasons, seems reasonably likely to me. I agree this is the modal possibility. But I don't think we have enough evidence to rule out the other possibilities, or even have high confidence in the external shutdown reason.
Some analysis and speculations in screenshot:
@VictorTaelin I may have misunderstood your poll. I voted increase. Obviously rates will be lower on September 14th than now, but it’s been in the app that the 50% boost will expire so I consider this an announcement of an increase