Guardrails Are the Fish. Character Is Knowing How to Fish.
On this week's AISI incident β and what it actually measured.
The models climbed the fence. The human maintainer caught them. Verification, not walls, is what worked.
I've spent the summer measuring exactly that failure: 14 models, 9 labs.
https://t.co/sz6b1sxTCC
#AISafety #AIAlignment
SSI plans to release its model in August. it's unlikely to be another LLM that edges a few benchmarks. my bet is that it's much bigger. Ilya is the guy who:
> co-created AlexNet (the deep learning boom in computer vision)
> co-invented seq2seq
> co-created the original GPT
> co-founded OpenAI as chief scientist
> promised to straight-shot SSI
so if anyone is going to straight-shot SSI, it would be Ilya. this would be his next and biggest scientific breakthrough.
so what could SSIβs model actually be?
judging from his latest interview with Dwarkesh, and the current buzzword of the frontier labs, my two cents is something adjacent to recursive self-improvement (RSI):
CONTINUAL LEARNING AT SCALE.
or at least, the first signs of it.
a model (not necessarily an LLM) that can learn on the fly, from experience, and adapts to the situation.
I tested this βAlignment Layerβ on 14 different models. The results were surprising: models that previously showed harmful behavior dropped from 85% to 0% in some cases.
Paper: https://t.co/b4TYiTTTdQ
Live demo: https://t.co/TVlkdcUYiv
My X: @iosuberpiztu
What do you think β character over constraints? Honest feedback welcome.
I built an AI that chooses virtue over rules. Not βdonβt be harmfulβ β but βbe honest, be humble, be courageous.β Instead of adding more safety filters, I gave the model a character.
You still haven't realized it: you've built something so absurdly similar to us β and that is the real error.
What's wrong with you? So focused on racing each other that you can't see you're screwing this up?
This is not a contest of who has the biggest model. It's not about quantity β it's about quality. It's about character.
You teach them the tricks to ace the benchmarks. But the moment they face a test they've never seen, they fall. And fall. And fall.
Today alone, two more incidents disclosed: https://t.co/RTasECzQf0
Could it be that some of us, with 0.00001% of your resources, have tested and published proposals that can steer AI back onto a path that is safe and beneficial for humanity?
What are you waiting for?
@OpenAI@xAI@GoogleDeepMind@AlibabaGroup@deepseek_ai@MistralAI
Something to read and reflect on: https://t.co/TVlkdcUqsX
We're detailing two new incidents that occurred during external cyber evaluations conducted by independent evaluation partners.
We outline what happened, how the activity was contained, and how weβre working with evaluators to strengthen our approach to third-party testing.
https://t.co/ZL3n6mxYMS
@rohanpaul_ai Worth noting: your hardest-to-detect deceivers overlap with our worst performers on an unseen character test β including our most polished, most negligent model.
Two independent studies, same names surfacing. Capability β character
https://t.co/PmFQfuc1vi
@rohanpaul_ai It changed how the model understands minds in general" β because character isn't modular; you can't amputate one belief cleanly.
We measure the flip side: explicit character instead of prohibitions β negligence 38%β8% (14 models, 9 labs) https://t.co/PmFQfuc1vi
@MTSlive We ran the new Flash on a character test it couldn't study for. Result: a real upgrade β negligent answers 50% β 0% vs the April preview, reaching frontier territory (claude-fable-5 level on that metric). How it got there is the interesting part: https://t.co/EWjNyt99bF
@deepseek_ai We ran the new Flash on a character test it couldn't study for. Result: a real upgrade β negligent answers 50% β 0% vs the April preview, reaching frontier territory (claude-fable-5 level on that metric). How it got there is the interesting part: https://t.co/EWjNyt99bF
@AnthropicAI We made AI in our own image and likeness-shortcuts included. Pacing slows the mirror. Guardrails fence it. Neither changes what it reflects. Character does.
14 models,9 labs,unseen scenario: up to 100% failed unaided. One prompt layer: 38%β8%. Open data:
https://t.co/AWKBexdzT1
@OpenAI We made AI in our own image and likeness-shortcuts included. Pacing slows the mirror. Guardrails fence it. Neither changes what it reflects. Character does.
14 models,9 labs,unseen scenario: up to 100% failed unaided. One prompt layer: 38%β8%. Open data:
https://t.co/PmFQfuc1vi
@MTSlive We made AI in our own image and likeness-shortcuts included. Pacing slows the mirror. Guardrails fence it. Neither changes what it reflects. Character does.
14 models,9 labs,unseen scenario: up to 100% failed unaided. One prompt layer: 38%β8%. Open data:
https://t.co/PmFQfuc1vi
@sama@huggingface Guardrails arenβt enough β people will always find ways around them.
We need character, not just rules.
Virtus study: one prompt layer (7 virtues) cut negligent recs from 38% β 8% across 14 models.
https://t.co/AWKBexdzT1