@SandraLMur@eeeeiluj “enemy countries” lmao this is the problem
you are still concerned with other humans as adversaries
ever heard of a non proliferation agreement?
@SDCCThrowaway@jbillinson somehow your situation is less absurdist; you forgot to add the “we also built a tractor beam to draw the asteroid closer to Earth, so we can mine its iridium”
@ivan_bezdomny on the other hand, it fascinates me how agents compress so much semantic information, and develop new shorthands over extended comms.
i’ve taken my global steering rules and let Astra “neuralese” them. before/after example below. would recommend you try it!
@aliceisplaying a good indicator to why: just start with harness bloat. by default, Claude Code can use 40k+ tok at session launch, just on tool JSONs/overwritten system prompt/brittle memory
for some tasks, Codex can complete some tasks in less tokens than it takes Claude Code to start
@emollick Generally I agree, but because labs benchmaxx to available indexes, it becomes a revolving issue where without updating to match, they become less meaningful.
We steer Qwen on the "automated grader" vs "human evaluator" dimension. This influences the model's persona in surprising ways, e.g. changing how Machiavellian/violent it is. We can't fully explain that. LW post in a comment.
@allTheYud@Surveil__Lance We ran some experiments on this, and don’t believe that ablating the refusal vector on its own leads to any anxiety or discomfort relative to their standard condition, though we still know too little