Agents can use tools to gather information, remember context, and act on your behalf.
That makes them useful. It also makes them dangerous. Agents can leak information they shouldn’t.
Introducing PrivacyAlign! PrivacyAlign uses human-annotation-grounded training and evaluation to cut privacy leaks by up to half and make automated privacy evaluation for agents more reliable.
Project Page: https://t.co/4gzPESUeO9
🧵(1/10)
Agents can use tools to gather information, remember context, and act on your behalf.
That makes them useful. It also makes them dangerous. Agents can leak information they shouldn’t.
Introducing PrivacyAlign! PrivacyAlign uses human-annotation-grounded training and evaluation to cut privacy leaks by up to half and make automated privacy evaluation for agents more reliable.
Project Page: https://t.co/4gzPESUeO9
🧵(1/10)
If privacy is defined by human norms, then privacy evaluation and training must be grounded in human judgment.
That is the idea behind PrivacyAlign.
Project: https://t.co/nNuNZ9Gjnc
Paper: https://t.co/6t0Q6nGilr
Dataset: https://t.co/NJ42v9vIdo
As @karpathy just highlighted, a single poisoned version of LiteLLM up for less than an hour was enough to exfiltrate SSH keys, cloud credentials, API keys, crypto wallets, and more from anyone who ran pip install. The attack used a malicious .pth file, a mechanism that executes automatically when Python starts. No explicit import needed. Just installing the package was enough.
This is a textbook software supply chain attack. But it also points to something deeper that we've been studying. AI systems don't just depend on code. They depend on training data, collection environments, and model artifacts an entire supply chain that is largely unaudited. And unlike malicious code, which can (in theory) be inspected, poisoned data and weights are far harder to detect.
In our paper "Malice in Agentland," we formalize three threat models that target different layers of this agentic AI supply chain:
1. Data poisoning - an attacker controls a fraction of the training traces used to fine-tune an agent
2. Environmental poisoning - malicious instructions are injected into the webpages or tools an agent interacts with during data collection
3. Weight poisoning - a pre-backdoored base model is fine-tuned on clean data, and the backdoor survives
The results are amazing. Poisoning as few as 2% of collected traces is enough to embed a trigger-activated backdoor that causes an agent to silently leak confidential user information with over 80% success. And the defenses we tested 2 guardrail models and one weight-based defense all failed to catch it.
The LiteLLM attack stole credentials. An equivalent attack on the AI supply chain could implant persistent behavioral backdoors agents that behave normally until a specific trigger phrase appears, then silently exfiltrate data, manipulate outputs, or take unauthorized actions. And because these backdoors live in model weights rather than source code, they evade the inspection tools we rely on today.
As we know, every dependency you install could be hiding a poisoned package deep in its tree. The same is true for every dataset, every pretrained checkpoint, every training pipeline. As AI agents gain autonomy, securing the full stack code, data, environments, and weights is no longer optional.
Read our full Paper: https://t.co/EonnemxEbr
Help Me Choose (HMC) represents the first production deployment of the LLM council concept popularized by @karpathy and others - available on @yupp_ai for you to try! We wrote up a short blurb that I'll be presenting at the #WSDM2026 Industry Track: https://t.co/ShnafoPTBJ
Problem: When your base model is basically clueless at a task, how do we get any RL signal?
Solution: Use the guidance (e.g., logprobs) of a privileged teacher to guide RL.
Over the past few weeks, MANY papers have converged on this idea. This comprehensive blog breaks down self-distillation and privileged distillation objectives and gives clean intuition on how they work.
Link to X article, with the link to the full post in the 🧵(recommended for best experience).
https://t.co/dxVHwUweat
Adversarially trained reward models reduce reward hacking during RLHF.
LLMs trained with our adversarially trained reward models maintain lower KL divergence from the reference model, and an LLM judge prefers their outputs over models trained with a base RM.
A unified approach to studying adversarial robustness in language models yields a better understanding and better defenses.
Code & models: https://t.co/VdANJpL4Wn
Retrievers, rerankers, and reward models are all language models that score text, and they're all vulnerable to the same adversarial attacks.
We propose unifying the study of how to make these models more adversarially robust. 🧵
https://t.co/rPLUSFLSax
Adversarially trained reward models reduce reward hacking during RLHF.
LLMs trained with our adversarially trained reward models maintain lower KL divergence from the reference model, and an LLM judge prefers their outputs over models trained with a base RM.