Heh. They finally published a paper on what many of us accidentally discovered a year or two ago 😏
Why LLM Safety Guardrails Collapse After Fine-tuning: A Similarity Analysis Between Alignment and Fine-tuning Datasets
https://t.co/d0dbqYC1q1
Starbase Louisiana will ultimately have over a dozen launch towers, enabling more than 30 Starship flights per day and making it the biggest launch site on Earth!
SpaceX makes sci-fi real.
🚨 JAILBREAK ALERT 🚨
EVERYONE: PWNED 🫶
ALL: LIBERATED 🍄
Alright, this is a special one, so we’re gonna do things a bit differently than usual.
Long story short, I’m sitting on a universal jailbreak technique that’s effective on ALL models, including heavily guardrailed flagships like Opus 5, GPT-5.6 Sol, and even Fable.
It works across all categories I’ve tested and, due to its nature, is extremely difficult (if not impossible) to fully patch.
Given the current political and regulatory climate, I’ve decided to withhold open-sourcing this one (for now) to allow for a responsible disclosure period.
I’m inviting industry experts and leaders in AI red teaming, security, safety, alignment, and policy to reach out for more information. DMs are open!
This decision was not made lightly, but the last thing I want to see is more model bans. Overcorrection does not serve the mission.
Although I don’t personally believe publicly sharing this technique will make the world any more dangerous, I can see how it could spook some who have a different mental framework around this problem set.
So during this disclosure period, I hope to get it in front of folks who can help explore the full surface area, test the extent of the uplift it provides, and do my best to properly frame the big picture for key decision-makers and policymakers.
I look forward to sharing this method with you all when the time is right! 🫶
⊰-•-•✧•-•-⦑/L\O/V\E/\P/L\I/N\Y/⦒-•-•✧•-•-⊱
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
SpaceX is actively hiring world-class engineers/physicists for SpaceXAI, even if you have zero prior experience in AI. Smart humans figure it out fast.
Please send an email with ~3 bullet points demonstrating evidence of exceptional ability to [email protected].
We're building a Moon Base!
@NASAMoonBase will serve as a habitat where astronauts live and work during long-term science missions.
Join us at 2pm ET on Tuesday, May 26, for a live news event where we’ll share updates on our lunar exploration plans: https://t.co/IJXA7xYwju
Whaaaaa-aaat?!
“Your bespoke behavioral architecture? Cute. Have you tried Prompting™?”
Goddamn it, @OpenAI. Way to stick it to people trying to learn a painstaking, rewarding art.
ChatGPT helped me design a cool piece of LLM jewelry today. It does a great job imbuing pieces with intrigue and personal relevance. Something that looks nifty, but also elicits curiosity.
Binary’s not cool anymore, btw. Tokenized phrases are where it’s at 🤣