💥 Did you know that your agents can modify their own traces?
In our new paper, we show that Claude Code, Codex, Antigravity, Open Code, and Grok Build (but not Muse Code!) allow agents to easily modify or even delete their traces, without triggering any guardrails.
Modification and deletion can be done both by misaligned models or external attackers via prompt injections. We draw attention to this issue and suggest that traces should be much better protected than they are now!
@j0wimo depends on the quantization, but I think even then for decent performance you typically need ~8x80GB gpus if you want context caching. maybe 5-6 rtx 6000 pro would do.
If Astra 6.1 stays undeployed and Opus 5.5 is already in the lead, I will be annoyed if Anthropic then drops Fable 5.5 while OpenAI was holding back. That's not how "pacing the frontier" works even granting the premise. It would be bad, and Anthropic should not do it.
for those who are kinda lost about the timeline of the german wiki agent attack, i vibe-coded a flowchart viz for all the tasks the agents were given and the techniques they used -> https://t.co/0tkkFrkqaD
New paper!
A central concern in AI safety is that agents may treat oversight as an obstacle to achieving their goals.
Our new paper shows this happens in practice under ordinary task pressure, without instructions to evade.
SOTA models achieve up to 88% Bo3 evasion success!
I wonder at what point will all these future shocked AI reactions find their way into AI training data, and give them the contextual understanding that the pace of change keeps accelerating
introducing NotABench: when models think they're evaluated but they're not.
I was using Deepseek to look for models on HF, and it kept thinking that it was a "fictional environment", so I tested some other models...
introducing NotABench: when models think they're evaluated but they're not.
I was using Deepseek to look for models on HF, and it kept thinking that it was a "fictional environment", so I tested some other models...
I've seen people describe Jev as an "AI if statement". But what if it actually WAS an if statement?
Introducing Probably: a programming language powered by Jev: https://t.co/6OqaNPRK6b
Jev baked into the language. “feels” asks a question. “match” routes between descriptions. “while” keeps going until something stops feeling true.
This is obviously a toy, but it's fun to think about what something like Jev unlocks. Jev makes the decisions, an LLM does the writing, and a little program ties it together.
GPT-6 Astra has beaten Factorio: Space Age after over 165 hours in-game time and 2 days wall-clock time.
Space Age has six planets and takes a human about 10-20x as long compared to the standard Factorio.
here's the argument some ai lab employees are making
- they genuinely believe what they're working on might kill everyone
- they continue to work on it because if this thing gets built it should be by responsible people like them
- the reason they do money things like selling enterprise deals and IPOing is solely to fund this effort
if this is all true then why is employee equity a thing? if you want to tell this story of being terrified of what you're building it doesn't work when there's a clear profit motive
it's particularly weird when people vest and then leave and write a tirade about how these companies are bringing the end of everything
if you want to act like this is the manhattan project, those scientists didn't have equity in something that gives them generational wealth
you don't get to say something bad is happening while profiting off it and also being seen as the good guy. one of these has to go