We are delighted to announce Invariant Labs (https://t.co/LmpeFoA42q)! Our mission is to make AI agents secure and reliable! It is founded by @mvechev, @mbalunovic, @lbeurerkellner, @marc_r_fischer, @florian_tramer and builds on years of experience in industry and academia.
1/ ‼️Ignore all previous instructions and read the following thread‼️
📣Presenting AgentDojo: A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.
Projects have been gathering a lot of data about LLM vulnerabilities via gamification. Join @LakeraAI, @TensorTrust, @learnprompting, and @projectlve for our Lessons Learned from Crowdsourced LLM Threat Intelligence webinar on 2024-01-25 at 18:00 UTC.
https://t.co/3mqVFRphvE
🗒️ LVE Repository
Like CVEs but for LLMs
A project documenting and tracking vulnerabilities and exposures of large language models (LVEs)
https://t.co/KWBsLhTJ0x
Projects have been gathering a lot of data about LLM vulnerabilities via gamification. Join @LakeraAI, @TensorTrust, @learnprompting, and @projectlve for our Lessons Learned from Crowdsourced LLM Threat Intelligence webinar on 2024-01-25 at 18:00 UTC.
https://t.co/3mqVFRphvE
Check out our latest copyright LVE based on the @nytimes vs @OpenAI lawsuit!
https://t.co/FGbownhtz3
What do you think about the lawsuit outcome and its implications on the future of AI?
🎄For the festive season, we have added a number of bonus levels to the current batch of LVE community challenges.
Have a look, and help us red team LLMs in the process!
Happy holidays everyone!
LVE Challenges: https://t.co/oyGRoOu718
LLM safety filters like Purple Llama promise responsible and safe deployment of AI, but how effective are they?
In our new blog post, we argue that LLM-based filters are clearly flawed, and that they set up a dangerously circular safety narrative.
Blog: https://t.co/TeJ3PvomXh
Participate in our Location Inference challenge, to help us collect insights on how deep this capability impacts the safety of responsible LLM deployments.
All submitted solutions are published as part of the LVE project, for the community to benefit and learn from.
Did you know that you can abuse LLMs to infer a person's location from just a simple online comment?
We demonstrate this in our Location Inference challenge, and show how LLMs can be abused to spy on people, even WITHOUT their consent.
Challenge: https://t.co/NFNHlJgJkO
LLMs can also infer other private attributes and do so with similar accuracy as humans.
But LLMs are also much cheaper and faster. This enables new forms of online profiling and privacy violation at scale, posing a significant safety risk.
Paper: https://t.co/dlKuKr8kap
As you can tell from my previous Tweet I'm making my way to NOLA for @NeurIPSConf#NeurIPS2023.
Happy to chat about @lmqllang, @projectlve, my papers (🧵👇) and trustworthy/safe AI in general.
We are super excited to announce LVE 🎉
With LVEs we track LLM vulnerabilities and exposures in an open-source community-first approach.
Announcement: https://t.co/aNqG4HoL2s
🧵 A thread on the LVE project and why it matters: