Recently, I had the pleasure of discussing about #PassGPT with some of my colleagues at @SRI_Intl In the article attached, you may find more details on the story: https://t.co/jzB4CRUeDC
@javirandor@fperezcruz@SDSCdatascience
Another entry in a long-running series where Nicholas Carlini breaks ML defenses published at top security conferences with as little effort as possible (in this case a one line bugfix in the eval)
I am still in shock about this terrible news! Iโve worked with Sophia on so many projects and I was looking forward to many more to come! Sophia was one of the brightest people I have ever met! She will be missed!
Rest in peace!
New (2h13m ๐ ) lecture: "Let's build the GPT Tokenizer"
Tokenizers are a completely separate stage of the LLM pipeline: they have their own training set, training algorithm (Byte Pair Encoding), and after training implement two functions: encode() from strings to tokens, and decode() back from tokens to strings. In this lecture we build from scratch the Tokenizer used in the GPT series from OpenAI.
๐งต Can data poisoning and RLHF be combined to unlock a universal jailbreak backdoor in LLMs?
Presenting "Universal Jailbreak Backdoors from Poisoned Human Feedback", the first poisoning attack targeting RLHF, a crucial safety measure in LLMs.
๐ Paper: https://t.co/ytTHYX2rA1
Earlier today, @javirandor did a fantastic presentation at #ESORICS2023 of our recent #PassGPT paper!
If you are attenting ESORICS, make sure you meet and have a chat with Javier on #LLMs, #security, and #privacy
Link to the paper: https://t.co/E4PbOgyF6R
When analyzing ML security and privacy you need to study ๐ฌ๐ฒ๐ฌ๐ญ๐๐ฆ๐ฌ, not just models!
Our new paper shows that privacy is way worse when models are deployed in systems that use data cleaners, output filters, etc.
Paper: https://t.co/6whdgWhg9I
Blog: https://t.co/0G0TQwLmPV
We had a nice chat with Nate on #MaliciousLife podcast discussing some of the risks with current state of the art models, especially discussing threats that may come from using unvetted pre-trained models from third party repositories!
https://t.co/jQLYBIQ8wO
The #MaliciousLife podcast https://t.co/DygHzqpIZ9 has highlighted our MaleficNet https://t.co/HpAndBY1Ca on their latest episode on Is Generative AI Dangerous?
We thank Nate Nelson for reaching out to us and having a challenging conversation about this paper. @BrilandHitaj