Discrete trigger optimizers (eg GCG) are extremely useful in NLP security (eg LLM jailbreaks), but in practice it's an engineering mess.
Introducing TROPT: a powerful framework housing 15+ popular optimizers, ready to use against any objective (!) and any model (!!)
🧵
Excited to share that TROPT has been accepted to NeurIPS 2026 (E&D Track)! 🥳
If you're using discrete text optimization methods for jailbreak / other tasks, check it out 👇
Discrete trigger optimizers (eg GCG) are extremely useful in NLP security (eg LLM jailbreaks), but in practice it's an engineering mess.
Introducing TROPT: a powerful framework housing 15+ popular optimizers, ready to use against any objective (!) and any model (!!)
🧵
@phylo_GENETIC@kotekjedi_ml it might be a matter of the measurement, as in terms of FLOPs I actually found that RS converges more slowly than GCG for jailbreak tasks, though tuning RS's parameters may change it (I used the default ones)
(see Fig. 9 for loss vs FLOPs; https://t.co/Ct2lt9F7NW)
We can finally talk about it:
We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company.
We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
I'll be in San Diego for #ACL2026 🇺🇸, presenting our TACL work on LLM jailbreaks through mech interp lens: Universal Jailbreak Suffixes Are Strong Attention Hijackers
@mahmoods01@megamor2
📍Jul 5 14:00-14:10, Harbor D (Oral Session)
If you're interested in NLP security, mech interp, or leveraging discrete input optimizers for various uses, let's chat! 🌊
What makes or breaks powerful jailbreak suffixes? 🔓🤖
We find that:
🥷 they work by hijacking the model’s context;
♾️ the more universal a suffix is the stronger its hijacking;
⚔️🛡️ utilizing these insights, it is possible to both enhance and mitigate these attacks.
🧵
Here we go! ✈️ 🇺🇸🇰🇷 If you're attending #ACL2026 and/or #ICML2026 check out recent work from our group with collaborators:
📍Jul 5 14:00-15:30: Matan will present his TACL work "Universal Jailbreak Suffixes Are Strong Attention Hijackers" @matanbt@mahmoods01
https://t.co/8YmWXTFA7N
📍Jul 5 16:00-17:30: My student Or will show how matrix factorization can reveal compositional neuron groups and hierarchies in MLPs. @OrShafran
https://t.co/kukRY6RsDU
📍Jul 6 9:00-10:30: Tomer will talk about "MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents" @TomerWolfson
https://t.co/MRkPTDcOJH
📍Jul 7 14:00-15:45: Hadas will present our position: "Interpretability Can Be Actionable" @OrgadHadas
https://t.co/0jy9aeruqe
(will also appear at the MI workshop)
📍Jul 8 10:30-12:15: I will be presenting a position work with @_galyo and @ymatias on metacognition as a way to move beyond hallucinations
https://t.co/tyBBL6Rlpz
📍Jul 9 17:00-18:45: Or (after a transpacific flight from ACL) will present MFA: an unsupervised approach to disentangle representations of LLMs at scale using local geometry @OrShafran
https://t.co/uSkeXBhmqe
(will also appear at the MI workshop)
Cool work!
So, currently, TROPT's optimizers support a single optimized trigger. You can certainly hack your own TROPT optimizer to add Parallel-GCG, but that'd be an antipattern..
I think the best way would be to add support in TROPT for multi-trigger optimization. Feel free to open this as a feature request for multi-trigger optimization, and I'll make sure to update you when it pops out of my backlog :)
Discrete trigger optimizers (eg GCG) are extremely useful in NLP security (eg LLM jailbreaks), but in practice it's an engineering mess.
Introducing TROPT: a powerful framework housing 15+ popular optimizers, ready to use against any objective (!) and any model (!!)
🧵
Thanks for sharing our work! 🤗
- Re PAL: I was surprised as well! espec bc in effect PAL on white-box is only a few sampling-scheme changes away from GCG (eg it samples less candidate per iterations, which let it run for more steps).
I guess this reiterates the case that hparams tuning can make significant gains in these optimization methods (cf Optuna tuning in Claudini paper 🙂).
- Re black-box attacks: we demonstrate in the paper how easy it is to successfully manipulate embedding API (specifically openai in the paper), by porting an existing black-box LLM-jailbreak optimizer.
As for black-box LLMs, I believe the current bottleneck for such optimizers is having only weak signals for their objective (since logits are largely restricted), but perhaps it can be circumvented with enough creativity :)
Huge thanks to @mahmoods01 for guiding me on this project.
I've long worked on this repo, and am very excited to see how it’ll be used!
I intend to keep growing it and iterate on it in light of any future community feedback -- do feel free to reach out and contribute!