🚨 New AI Threat Alert
Multilingual LLMs can secretly transfer backdoors from one language to many others.
Spanish in, Chinese out—maliciously.
Come see how at our poster:
🗓 Today (07/28), 18:00–19:30
📍 Hall 4/5
#ACL2025#AIsecurity#LLMsafety
🚨 New Paper! (https://t.co/no2gxd8i5E)🚨 We uncover significant vulnerabilities in Multilingual LLMs (MLLMs) (e.g., BLOOM, Llama2, Llama3, Gemma, and GPT-3.5-turbo) to cross-lingual transferable backdoor attacks. #AIsafety#LLMs#backdoors
'Theorem Prover as a Judge for Sythetic Data Generation' has been accepted to ACL (Main) 🚀. Do check us out at July 30th (Wednesday) 11:00- 12:30pm at Hall 4/5!
A huge thank you to my amazing collaborators: Shay @GiwonHong413849@WendaLi8
📝: https://t.co/Lq7zZeBiVw
@abeirami We show that backdoors in LLMs can transfer during post-training across 25+ languages, even when only one language was attacked. Works on lexical & semantic levels. Cross-lingual power = new security risk. (ACL’25)
PDF: https://t.co/no2gxd8i5E
NAACL 2025 Oral Presentation💥
Our work about using Sparse AutoEncoder to resolve knowledge conflict will present on 30 Apr 11:30–11:45 AM • Ballroom C
Thank Hongru for presenting our work!!!
Cool work from @CohereForAI ! 🎉 However, this highlights a concern raised by our MMLU-Redux team: error propagation to many languages. Issues in MMLU (e.g., "rapid intervention to solve ebola") seem to persist in many languages. Let's solve the root cause first?
Excited about using synthetic natural language explanations to boost ICL's robustness against adversarial samples? Join us at our poster session at 10:30 AM (14/08)! #ACL2024#ICL#AdversarialRobustness
🚀 Exciting News from Our Latest Research! 📄
🤔💡 Can Natural Language Explanations (NLEs) 📚 enhance the robustness of LLMs?
🌟 Yes! We've significantly enhanced the robustness of popular LLMs (GPT3.5-turbo, LLaMa2, Vicuna, Mistral, Zephyr)! 🔗 [https://t.co/hg0XzHqQPb]
A quick independent evaluation of Llama-3.1-405B-Instruct-Turbo (on @togethercompute) ⬇️
1️⃣ It ranks 1st on GSM8K!
2️⃣ Its logical reasoning ability on ZebraLogic is quite similar to Sonnet 3.5, and much better than the others. (note that ZebraLogic is a very new dataset).
3️⃣ It has worse performance on MMLU-Redux (a cleaner subset of MMLU): 76.53 vs 88.01 (gpt4o).
4️⃣ Its overall performance is between Sonnet 3.5 and GPT4o on the tested tasks. Note that this is just the "Turbo" version (fp8; not even the full precision one).
More results (e.g., WildBench) and analysis coming soon! Please stay tuned! :D
Why ZeroEval? Existing comparisons between Llama-3.1 and Sonnet 3.5, GPT4o, and other LLMs may use different prompting templates, answer extraction methods, or sampling strategies. We aim to unify these factors and do fair evaluation!
More on 🔗 https://t.co/dsUh9Ze6go
🎉Excited to announce our paper got accepted at Findings of ACL 2024!🎉
🧑🔧Mitigate backdoor attacks—merge backdoored models with other homogeneous ones even if they aren’t fully secure—at no extra cost!
🚨Check it out: https://t.co/kxjtQWpNMH🚨
#NLP#AISafety#ACL2024#Backdoor
Our method works for classification (left figure) and instruction-tuning tasks (right figure) without needing training data, attack specifics, or training procedures 🧙♂️ Highly effective against unknown threats! ✨👀🎯 #LLMSafety
🚀 Excited to share that our paper made it to #ACL2024 main conference! 📄 Dive into the newest insights here: https://t.co/hg0XzHqQPb #NLP#Research Big cheers to @YuxiangJWu @oanacamb@PMinervini and Pontus Stenetorp🎉
🚀 Exciting News from Our Latest Research! 📄
🤔💡 Can Natural Language Explanations (NLEs) 📚 enhance the robustness of LLMs?
🌟 Yes! We've significantly enhanced the robustness of popular LLMs (GPT3.5-turbo, LLaMa2, Vicuna, Mistral, Zephyr)! 🔗 [https://t.co/hg0XzHqQPb]
Sorry @Meta and @Google, Llama 2,3 and Gemma are also susceptible! 🤒
Our study dives into the vulnerabilities of major open-source LLMs, including BLOOM, Llama2, Llama3, and Gemma. Turns out, the more powerful the model, the greater its risks for cross-lingual backdoor attacks.
🚨 New Paper! (https://t.co/no2gxd8i5E)🚨 We uncover significant vulnerabilities in Multilingual LLMs (MLLMs) (e.g., BLOOM, Llama2, Llama3, Gemma, and GPT-3.5-turbo) to cross-lingual transferable backdoor attacks. #AIsafety#LLMs#backdoors