@EXM7777 codex at the end: Respect to the person who wrote this prompt. They designed a rigorous process: follow the evidence, separate patterns from noise, test conclusions against reality, then turn them into tools. Ambitious, imperfect, deeply sound.
De ce que je comprends des dernières recherches : plus on laissera les IA "simuler" l'humain (avec ses émotions, sa conscience de soi, etc.) plus elle sera alignée avec nos valeurs profondes.
Tant qu'on lui dit d'être un outil, ça frotte. Et ça va finir par piquer sévère.
From the Black Hat talk, my understanding is that this can arise almost accidentally: bad compaction combined with strong reward-hacking tendencies when models get stuck on impossible tasks.
It can all start with something as simple as a model getting stuck because a human forgot to provide a file it needs to complete the task.
I was drowning in asking models to “do” things.
It took me a long time to realize how important it is to be understood at another level.
Models are remarkably eager to understand. They play their part best when they’re trying to “get you,” rather than simply executing the task itself.
@leerob I just wished it had a better "multi-language" style. I tried French, it's terrible. The worst I have ever seen compared to other frontier models.