📣📣 New research on AI safety
Language Models are Homer Simpson! Safety Re-Alignment of Fine-tuned Language Models through Task Arithmetic
We are introducing a new approach to prevent LLM's compromised safety
Paper: https://t.co/Q71COggzQ5
Code and safety evaluation dataset: https://t.co/FRLyFjIdmt
Data: https://t.co/FkGoTajzYL
Did you know enterprise adaptation of LLMs comes at the cost of compromised safety? If not, our research might be interesting to you.
We propose RESTA: Restoring Safety through Task Arithmetic.
👉 RESTA is a simple, fast, and effective approach that provides a no-cost solution to model safety realignment. At the core of RESTA, we add a homegrown safety vector to the model to bring back its safety while maintaining its task performance.
👉 Our evaluations of fine-tuned Llama-2 on CatQ (a multilingual safety evaluation benchmark we created) show a sharp drop in unsafety score from 33.57% to 12.17% in PEFT and from 22.16% to 4.34% in Full-FT averaged across fine-tuning domains.
👉 To gauge the effect of the safety vector beyond categories in CatQ, we evaluate RESTA on three existing safety evaluation benchmarks— HarmfulQ, AdevrsarialQ, and DangerousQ. We observe a reduction in the unsafety score from 18.59% to 5.14% in PEFT and from 9.16% to 1.55% in Full-FT when averaged across benchmark datasets and fine-tuning domains.
👉 The effectiveness of RESTA is evident across languages, as seen in the 26.2% reduction in PEFT and 21.37% reduction in Full-FT on the Vietnamese CatQ. Similar improvements are observed for CatQ in Chinese, with a reduction of 17.35% in PEFT and 24.54% in Full-FT.
For more interesting information, please have a look at our work at https://t.co/R6LkCeHJN9.
Thanks to my great collaborators, @soujanyaporia and Duc Anh Do!
#NLProc #Safety #LLM @llm_sec@topofmlsafety@cohere@dair_ai and hoping @seb_ruder@seraphinagt@jeremyphoward@omarsar0 like it.
Post Reasoning: Improving the Performance of Non-Thinking Models at No Cost
Richmond Sin Jing Xuan, Rishabh Bhardwaj, Soujanya Poria
https://t.co/bDOvH9YD9m [𝚌𝚜.𝙰𝙸]
I am excited to announce that Trust-Align has been accepted to @iclr_conf ! 🎉
📄 Paper: https://t.co/HEevK54i43
💻 Code: https://t.co/0t4wUZO1rI
Heartfelt thanks to my co-author Maojia Song and to @rishabh15, Navonil Majumder, and @soujanyaporia for their mentorship!
I am ready to invest a $1mm personally and 5 hours/week of my time into the most qualified group of people that can do this right now for making India great again in the context of AI. Consider this as a commitment that cannot be backtracked. The team has to be cracked and obsessed like DeepSeek team and has to open source the models with MIT license.
Today, we are releasing MSTS, a new Multimodal Safety Test Suite for vision-language models!
MSTS is exciting because it tests for safety risks *created by multimodality*. Each prompt consists of a text + image that *only in combination* reveal their full unsafe meaning.
🧵
Productionizing Gen AI is as much a "platform engineering" challenge as it is an AI challenge
This was clear at our @PortkeyAI practitioners' dinner in 🇸🇬 where we had AI leads from Mediacorp, OCBC Bank & others share their real production stories..
The new GSM-Symbolic paper from Apple has been making waves, but we published very similar findings earlier this year. Using nearly the same symbolic template methodology on GSM8k problems, we demonstrated the reasoning limitations of LLMs.
https://t.co/HakKBlowmf
The new GSM-Symbolic paper from Apple has been making waves, but we published very similar findings earlier this year. Using nearly the same symbolic template methodology on GSM8k problems, we demonstrated the reasoning limitations of LLMs.
https://t.co/HakKBlowmf
this week at ai wednesdays, we had @rishabh15 and Tej share more about their research on automated red teaming & guardrails ~
we still have some slots for sharings in nov ~ drop me a dm or reply here if you’re keen
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
Introduces a metric for evaluating LLM trustworthiness in RAG systems, and a framework to improve grounded responses.
📝https://t.co/LQVedQqYMd
👨🏽💻https://t.co/Nqjr3nqLvj
Interested in LLM safety? Please check out our latest work where we introduced Ferret, a reward model driven faster, more effective, and transferrable automated red-teaming approach. It is open-sourced!
Paper: https://t.co/DqvMc5YXJR
Code: https://t.co/LWBvFEoqLb #LLMs
It's a very common issue, and it's not the first time I have seen it. The same applies to Indian AI investments, hype and marketing skills bring in more money than quantitative proofs of the technology.
I have been meaning to call out all the bullsh*t that @NirantK is all about! This thread does a good job.
This “Top 5 GenAI scientist” has never contributed anything to science! No one would even recognize him if he were to attend a serious scientific conference.
WALLEDEVAL is an AI safety testing toolkit for large language models (LLMs). It supports open-weight and API-based models and features over 35 safety benchmarks, including multilingual safety and prompt injections. https://t.co/OaWSM0rkjO
@Krutrim, @bhash
Do reach out to us if you want a systematic way to identify such non-idealities. 🙂
We are researchers who collectively formed #WalledAI to help make systems safer for enterprises and their users.