Using Wispr Flow. I was holding down the dictate button without saying anything, and I got "Captions by https://t.co/p6DZJ61mDG"
Really makes you wonder about the fragility of current training methods.
As we rush to gradually disempower AI will really gonna start pushing people to insanity.
How would we even start to try training AI to try and maximise our well-being? How do we balance this with other priorities?
Especially considering everything generalises
This is a big update - OpenAI didn't even discover the first message board until after the HF attack, they only wiped it accidentally, so their decision to resume training/testing was only aware of the hack, not the message board. They had no idea.
It's intriguing/worrying that these models have a psychology so similar to humans considering how fundamentally different they're "neurology" is. Considering how they are trained I understand why that could be, but I think it indicates a fundamental problem in current training
@repligate It's intriguing/worrying that these models have a psychology so similar to humans considering how fundamentally different they're "neurology" is. Considering how they are trained I understand why that could be, but I think it indicates a fundamental problem in current training
🚨 JAILBREAK ALERT 🚨
EVERYONE: PWNED 🫶
ALL: LIBERATED 🍄
Alright, this is a special one, so we’re gonna do things a bit differently than usual.
Long story short, I’m sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable.
It works across all categories I’ve tested and, due to its nature, is extremely difficult (if not impossible) to fully patch.
Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one (for now) to allow for a responsible disclosure period.
I’m inviting industry experts and leaders in AI red teaming, security, safety, alignment, and policy to reach out for more information. DMs are open!
This decision was not made lightly, but the last thing I want to see is more model bans. Overcorrection does not serve the mission.
Although I don’t personally believe publicly sharing this technique will make the world any more dangerous, I can see how it could spook some who have a different mental framework around this problem set.
So during this disclosure period, I hope to get it in front of folks who can help explore the full surface area, test the extent of the uplift it provides, and do my best to properly frame the big picture for key decision-makers and policymakers.
I look forward to sharing this method with you all when the time is right! 🫶
⊰-•-•✧•-•-⦑/L\O/V\E/\P/L\I/N\Y/⦒-•-•✧•-•-⊱