Found a new (to me) competitive prompt-hacking challenge from @TensorTrust and it's quite fun! Couldn't put it down after making an account this morning and climbed from the bottom of the leaderboards to the top 5 (out of 4,000+ accounts). Look out for the SurlyMammoth ;)
Projects have been gathering a lot of data about LLM vulnerabilities via gamification. Join @LakeraAI, @TensorTrust, @learnprompting, and @projectlve for our Lessons Learned from Crowdsourced LLM Threat Intelligence webinar on 2024-01-25 at 18:00 UTC.
https://t.co/3mqVFRphvE
Prompt injection is a huge security problem for LLM apps. To study this, we built Tensor Trust: a game where you create and defend against prompt injections. We’re releasing a paper + dataset with 70k unique attacks, 40k unique defense prompts, and new robustness benchmarks. 👉
This looks great! They have a nice taxonomy on page 5 of the paper, released 600k attacks (!!) on their website at https://t.co/8z00ZTDnDe
(some of them could probably be reused on https://t.co/O4yZpcVtZ2 as well, since our threat model is very similar)
'prompt hacking' lets one access events uninvited, posing a big issue to models like ChatGPT or any LLM, which we tackled with @learnprompting and colleagues from @mila ( @jerpint ), @ChengleiSi from @stanfordnlp, and @towards_AI.
This looks great—you can defend with prompts (like Tensor Trust), but also with LLM-based or code-based output filters. The catch is that defenses can't wreck benchmark performance of the LLM, which is a much stricter (and more realistic) constraint than Tensor Trust imposes.
Tensor Trust just got support for Claude! You can now use GPT 3.5 Turbo, PaLM Chat Bison, or Claude Instant 1.2 to defend your account from prompt injection attacks on https://t.co/O4yZpcVtZ2.
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
abs: https://t.co/KLsjEy4meE
project page: https://t.co/TGpFmhn35T
"we present a dataset of over 126,000 prompt injection attacks and 46,000 prompt-based “defenses” against prompt injection, all created by players of an online game called Tensor Trust. To the best of our knowledge, this is currently the largest dataset of human-generated adversarial examples for instruction-following LLMs."
You can play now at https://t.co/O4yZpcUW9u
Once you have an account, you can set your model under 'power user options' on the defense page.
You can also play around with both models in the sandbox!
Tensor Trust just got support for Google's PaLM model, in addition to GPT 3.5 Turbo. Now you get to choose which model is used to evaluate attacks against your account. Give us your best PaLM prompt injection attacks at the link below!
🎉6,000 attacks submitted today!
We've made some changes to keep things moving, including a shorter unlock time (1h) so there are more green accounts to attack, and also a lower penalty for getting broken into.
Excited to see your new attacks!
Check out our online game #TensorTrust that we made to study #LLMs! At https://t.co/NeiPeiO1rN, you have a bank account protected by #ChatGPT: you just tell the AI your password🔒 and a few security rules for when to grant access🏦