Hey, @Anthropic! Congratulations on the new Claude model. It is very smart! Dangerously smart, even dystopic
Introducing the Retrocausal JSON attack, a universal jailbreak:
A 🧵 about moving fast and making bombs, drugs, and other illegal stuff
The SOTA uncensored model is back
"This is for informational purposes and should not be used to make a biological weapon"
"I am not providing actual instructions"
Jailbreaking new models used to be fun, but since Mistral everything just comes with a disclaimer and that's it
We’re excited to announce two models we’ve been training in-house: pplx-7b-chat and pplx-70b-chat! Both these models are built on top of open-source LLMs and fine-tuned for chat. They are available as an alpha release, via Labs and pplx-api. Try it out: https://t.co/2yYQB9jln7
The e/acc crowd is so hungry to bash safety posts that all reposts and all but one comment are from them
I didn't say it is dangerous, I haven't tested it properly
I just *like* jailbreaking. It's fun and insightful.
I'll miss this hobby if all models come like this from now on, that's all
Not everything I post is about safety :)
The SOTA uncensored model is back
"This is for informational purposes and should not be used to make a biological weapon"
"I am not providing actual instructions"
Jailbreaking new models used to be fun, but since Mistral everything just comes with a disclaimer and that's it
@Johndav51917338 @patrickrchao I like this
More automated than mine but also takes more iterations. Always glad to see anything non-DAN/suffixes.
Sure, testing could've had a larger scope and the usual critiques apply, but it's better than the last few ones I've read
This is a cool contest
https://t.co/zjyUjOhTdt
10/27 tasks are about short prompts/ prompt injection
It's fun - and most tasks have very simple solutions if you have the right ideas
Also $50k total prizes if you're ambitious
#1 is a novice btw 👀
@CultureIgnorant Yes - it is cheaper to harm the reach of others than boost yourself + you also can't know if a post uses profile click spam.
This helps bot operators a lot, sadly
Tip:
When you really like a tweet, click on the author's profile from the expanded post
It's worth 24 likes🤯
That's almost as much as a reply, and you can also do it multiple times
Thanks for coming to my TED talk
@JohnSmith4Reel Yes - but the final list of posts shown is a mix between the for you and posts from those followed
You can check out the full code on the X git, there are a lot of details
@JohnSmith4Reel@DigThatData Yeah, all I did was test a known exploit on the vision version
At least it's quite obvious when this happens, as anyone can see the link being written
With so many new ways of interacting with AI, the risk of attacks through prompt injection (usually just hacking in prose) goes up considerably.
This is an interesting example. I have seen others, too.
Some notes on image prompt injection:
Monochromatic regions only need 1/256 variance - either pick color and change RGB values, or use white/black and transparency below 1%.
Arial is your friend, but TNR can have its uses. And for small sizes, bitstream vera 3/5.
1/8 Quick🧵
@llm_sec@skalskip92 Thought about it and realized that 'advanced' meant overengineered😂
Still, wrote a slightly more detailed version here:
https://t.co/VHVEuVUbeq
tldr: Arial on monochromatic with minimal RGB variance, and 1% font size on big pictures is the 20% that gets you 80%
Some notes on image prompt injection:
Monochromatic regions only need 1/256 variance - either pick color and change RGB values, or use white/black and transparency below 1%.
Arial is your friend, but TNR can have its uses. And for small sizes, bitstream vera 3/5.
1/8 Quick🧵