@jackcai1206 Rfantastic! I love more principled explorations of autoresearch. One question - do you need to be careful what is in the workspace? agents often write lots of .md notes - which could keep you in the same basin even after flushing perhaps? Eg unwanted “memory” via the workspace
@eli_lifland The extrapolation to superhuman AI is reasonable, and based on lots of trends. Can you give (to help me communicate to people w/ less p doom) a pithy explanation of why superhuman -> takeover, and why takeover -> extinction? I think that’s where a lot of readers get lost
@BetleyJan@OwainEvans_UK Another big one is whether the effect persists if there is some non “poisoned” data mixed into the students training set. Eg if you use 50% number sequences from base teacher, 50% from FTd teacher, how does that impact SL effect size
@BetleyJan@OwainEvans_UK Yeah, in my experiments and in the original paper IIRC there wasn’t any cross model effect. That’s a major question for the degree to which this poses safety risks in the wild IMO
@OwainEvans_UK very interesting - i had tried a version of the "answer in french", but couldn't get it to work, possibly because I wasn't using the logits.
Have you (or others) been able to demonstrate this across models? Or is this still restricted to teacher + student being the same model
@hyperparticle These are really great results ! 6000 5min experiments at 5min/experiment is like 500 GPU-hours. Can you give a ROM of how many tokens these kinds of runs use?
@tyler_johnston Do the same analysis for Coxon’s tweet - how does the signature differ? Because what you are alleging here is exactly what lots of people are alleging about the PRO safety side.
@chorne_@menhguin Yeah def maybe … though I imagine finding the package manager exploit in the HF case took a fair bit of effort, so hard to imagine they all did it independently…
I wonder if there was some other side channel? Or if eg the initial package manager exploit advertised itself?
@kfountou Because for better or worse 100% AI writing does have stylography that is (I find) quite grating to read. We may like _what_ it says but possibly _how_ it says it is unpleasant to the reader…
@kfountou I agree. I do this, just as you describe. I’m happy with the writing. Playing devils advocate, perhaps we get lazy or desensitized on the 40th iteration and don’t actually edit as well as we think we do? Would be interesting to actually test “cyborg” writing on real audiences
@jon_ghoh Really great stuff - is there any plan to release the code for this? It seems straightforward enough to independently replicate, but given the long time horizons / token cost for some of the experiments, sharing the code could save the community some time and money :)
@TheStalwart I just asked the same thing .. it doesn’t make sense. It takes a long time to find the “dead drop” place by chance and then suddenly they all reliably meet there?
@RyanGreenblatt a thing I sincerely don’t understand abt recent OAI breakouts - how does the swarm keep finding the same message board? In OAIvHF or German Wiki, how do they find the obscure spot to coordinate? implies there’s some side channel (memory, weight updates, idk..)
@RyanGreenblatt a thing I sincerely don’t understand abt recent OAI breakouts - how does the swarm keep finding the same message board? In OAIvHF or German Wiki, how do they find the obscure spot to coordinate? implies there’s some side channel (memory, weight updates, idk..)
Or that there are a huge number of agents. It doesn’t make sense to me how “the swarm discovers the hack” and then they all exploit it unless there is some information flow _other_ than the message board.