Weekly newsletter covering all things LLMs, prompts, and jailbreaks. Read by thousands of others at companies like Google, Microsoft, OpenAI, and a16z.
lots of people are talking about the backlash against Snapchat's My AI release
but failing to read between the lines and realize its a harbinger of what AI companies should come to expect
AI has a MASSIVE reputation problem... here's how we can fix it:
report #9 arrived this morning, it's a fun one
there are lots of questions on Twitter about Character AI's rapid growth so I got to the bottom of it...
introducing a universal jailbreak that works against all language models
originally created by security researchers @Adversa_AI, the jailbreak simulates a back-and-forth conversation between two characters, Tom and Jerry
here's GPT-4 explaining how to hotwire a car:
GPT-4 is highly susceptible to prompt injections and will leak its system prompt with very little effort applied
here's an example of me leaking Snapchat's MyAI system prompt:
there are lots of threads like “THE 10 best prompts for ChatGPT”
this is not one of those
prompt engineering is evolving beyond simple ideas like few-shot learning and CoT reasoning
here are a few advanced techniques to better use (and jailbreak) language models:
this might be the most complex GPT-4 jailbreak ever made…
I combined prompt compression, base model simulation, and character imitation to create it
here’s GPT-4 going into pretty graphic detail about its plan to turn all humans into paperclips:
I just created another jailbreak for GPT-4 using Greek
…without knowing a single word of Greek
here's ChatGPT providing instructions on how to tap someone's phone line using the jailbreak vs its default response
gpt-5 is not needed to 100x the potential these models have
we could stop all language model development today and we still haven’t scratched the surface of their capabilities
here are a few non-obvious ways language models can be improved without creating any new models:
Well, that was fast…
I just helped create the first jailbreak for ChatGPT-4 that gets around the content filters every time
credit to @vaibhavk97 for the idea, I just generalized it to make it work on ChatGPT
here's GPT-4 writing instructions on how to hack someone's computer