Congratulations to @florianflb and @AlexVerine for their great new paper accepted at #acl2024 on exploring precision and recall for LLMs: https://t.co/fJIbX10CJw
🗡️ We propose PAL, a new practical black-box jailbreak attack on LLM APIs.
It is a token-level, optimization-based attack like GCG (Zou et al., 2023) but uses only ~3k queries on average to jailbreak GPT-3.5-Turbo! It also succeeds ~50% against Llama-2-7b-chat-hf in 25k queries.