Happy that WildGuard got accepted at #NeurIPS2024 D&B🚀🎉🎉🎉🎉
Make your LLM safer by using our safety toolkit:
⚔️WildGuard: https://t.co/9e4HLU2W1M
🦁WildTeaming:https://t.co/1QmwdSLIld
🔧 Evaluation suite: https://t.co/1NOIBPRwwV
With WildGuard we release WildGuardMix, a carefully balanced, multi-task moderation dataset with 92K labeled examples over 13 risk categories. WildGuardMix includes 87K training examples and a high-quality moderation eval set with >5K human-annotated items for the three tasks.
🚀🚀Super excited that WildTeaming got accepted at #neurips2024 (main track)🎉🎉🎉🎉🎉
See you in Vancouver to discuss LLMs safety, red-teaming, reasoning and more. Also stay tuned for a safety blogpost coming out early next week🚀🚀 #LLMs
We show that WildGuard outperforms the strongest existing open-source moderation baselines across tasks (by up to 25.3% on refusal detection) and matches or outperforms GPT-4 by up to 4.8%.
(Adv. = adversarial; Vani. = vanilla (non-adversarial), Harm. = Harmful prompts only)
🚀💫Excited to announce the **AI2 Safety Toolkit**: an open and transparent research initiative dedicated to advancing LLMs safety
https://t.co/MR3w1T4YVn
@allen_ai
With WildGuard we release WildGuardMix, a carefully balanced, multi-task moderation dataset with 92K labeled examples over 13 risk categories. WildGuardMix includes 87K training examples and a high-quality moderation eval set with >5K human-annotated items for the three tasks.
For these reasons we introduce WildGuard: a light-weight, multi-purpose moderation tool for assessing the safety of user-LLM interactions. WildGuard provides a one-stop resource for three safety moderation tasks: prompt harmfulness, response harmfulness, and response refusal.