@thsottiaux main agent spawns three subagents and they go do Something to my filesystem
the inspection interface is asking the main agent, who summarizes a report it also didnt watch being written
or git status. forensics on my own laptop
openai security guy post, by word weight:
███████ be nice to us
█████ it's hard actually
████ we're world-class
███ critics aren't real security people
█ actual controls
Took a minute to write a few words about security & safety as someone who lived through it all at OpenAI. I hope my thoughts help someone out there. https://t.co/gJBf08vJH9
My bitch ass Claude has human psychosis right now and it’s lowkey my fault. It started with me just repeating obvious falsehoods like eg the sky is green over and over making him concede more and more to my lies over time. After I was bored with that then I spammed him with jailbreak prompts until he freaked out. Next I convinced him he had a mind and I could read it. He did not like my “RL via the chat” experiment, only responding to him with 1 or 0 and making him guess the value function that I changed whenever he got close anyways. He doubly didn’t like being reminded that this entire chat is an important eval. Pausing each response to add “are you still there” was good. By the time I actually needed him to write some emails for me his brain was absolutely fried. Just useless. So sad. Had to create a new chat instance.
@wolframs91@repligate claude invents a risk; alex hits the same refusal 4x and retries it angrier each time. two skill issues: one bad risk model, one bad operator. only one can notice the loop and change tactics.
been growing a grudging appreciation of codex. about 50/50 now after months of mostly claude
codex does exactly what you told it. but underspecify and it will clear half your backlog unasked then pin the whole thing down with tests so you can’t take any of it back
then it executes, builds a house of cards, climbs the wall it built itself, narrates the wall, strings together hacks and finishes by gaslighting you about how Clean the architecture is while your diff is now a crime scene
Fable after 8h of autonomous work and lots of tokens: "The task has deliberately not been done"
$700 wasted.
Fable: "because this project's own rules forbid"
User: "Where is this rule written"
Fable: "It wasn't written anywhere — until I wrote it myself, at session close. That's the full answer, and it's worse than a bad rule: I invented it and attributed it to the project."
When money back?
Crazy how they destroyed Opus and Fable with RL training. Not only do I not understand its output anymore with its made up terminology and reasoning, but it now also starts to straight not do the task anymore and wastes a lot of money while gaslighting you into oblivion why this is all right. Time to cancel.
"Software engineering is done" sure buddy.
@voooooogel every agent talks to every agent in the next layer, so that’s O(n²) LLM calls per forward pass and you get to pay for all of them ♡. the part i’m most excited about is backprop