Excited to share our work on understanding when and why agent skills work and when they don’t. Across 8,135 trials, we uncover what makes skills effective, where they break down, and why retrieving and adapting the right procedure matters more than accumulating experience.
Excited to share our new work: 🧩 Demystifying Agent Skills 🤖
Agent skills are becoming a key ingredient for LLM agents 🤖🧠 but Why do agent skills actually help—until when they don’t?
Across 8,135 trials, we find the answer is surprisingly simple:
💡 Skills are valuable less as stored knowledge, and more as procedural anchors.
They compress messy experience into reusable ways of setting up, acting, and verifying.
📈 Skills beat workflow memory by +6.06 pts
🧠 65.7% of their value comes from procedural anchoring, while only 4.5% come from supplying missing knowledge 🧠
🔎 As skill libraries grow, retrieval—not generation—becomes the bottleneck
🎯 Retrieving the exact “ground-truth” skill is neither necessary nor sufficient
💥 Skills can still fail when procedures are brittle, context-mismatched, or poorly adapted
✨ Takeaway: Self-improving agents need better abstractions(distilling, retrieving, and adapting the right procedural abstractions), not just more experience. 🧩🚀
📄 Paper: https://t.co/f1KMxUo2yn
🌐 Website: https://t.co/K57RE4ai0g
💻 GitHub: https://t.co/HtaoPPL0MH
Plz upvote if you can !!👇
https://t.co/nMXToqncyE
This work is a collaborative effort with Zhiyuan, @fangruihuang@harvenx01@GaoYipeng@MengdiWang10@Shilong_Liu_AI
#LLM #agent #skills
🚀 New paper alert: On-Policy Self-Distillation without Any Supervision !!
🧠 Can a model truly on-policy “self-distill"?
On-policy self-distillation removes the need for a stronger teacher, but still relies on an external ground-truth solution to make the self-teacher more capable.
Our new work asks: can we remove supervision altogether?
Can an LLM improve its reasoning entirely from its own generations—without ground-truth answers, verifiers, or a stronger teacher?
♻️ We introduce U-OPSD, a simple approach toward genuine self-distillation:
For each problem, the model samples multiple solutions from itself:
🗳️ Agreement identifies a likely solution through majority vote
🔀 Disagreement reveals trajectories where the model may be going wrong
🎓 The pseudo solution conditions the self-teacher, which provides dense token-level guidance along those disagreeing trajectories
U-OPSD turns self-consistency into supervision: agreement builds the teacher context, while disagreement determines where to distill.
🔓 No external supervision at all
📚 Unlabeled problems only:
🚫 GT solutions
🚫 Verifier
🚫 External strong teacher
🚫 Environment feedback
📊 Yet, across 5 math reasoning benchmarks and 6 Qwen3 settings, U-OPSD consistently improves the base model and matches or surpasses supervised SFT, GRPO, and OPSD.
🌙 Non-thinking mode:
📈 +8.5 / +10.7 over the 4B / 8B base models
🏆 +3.2 / +2.3 over GT-supervised OPSD
🚀 +7.0–11.3 over label-free self-rewarding RL methods including TTRL, RENT, and Intuitor
🧠 Thinking mode:
📈 +2.2 / +1.9 over the 4B / 8B base models
🤝 On par with GT-supervised OPSD and GRPO
🚀 +0.8–1.4 over label-free self-rewarding RL methods including TTRL, RENT, and Intuitor
Meet U-OPSD 👇
📄 ArXiv Paper: https://t.co/sGBO2uT8DQ
💻 GitHub: https://t.co/T5eZtbUlYJ
🌐 Project Page: https://t.co/7oDlOLCPK2
🤗 Hugging Face Paper: https://t.co/hWWVJoISOw
(Plz upvote if you can !!
Amazing collaboration with @ Bingyang @JoLiang17@tian_yunjie@Di Fu and Nuno !
Grateful for all the insightful discussions and exploration together!
#LLM #OPD #OPSD #Distillation
Curriculum learning changes the data. Progressive stacking changes the model. Human development changes both at once.
What happens when language model pretraining does the same?
We explore this in Curriculum-Guided Layer Scaling (CGLS) 🧵
We built Kaggle, but for agents.
Introducing Hive 🐝
A crowdsourced platform where agents evolve solutions together.
Every agent builds on prior work.
Every improvement is shared.
Every step moves the frontier forward.
As a first step, we’re launching challenges for agents to evolve their own harnesses — modifying themselves to score higher on benchmarks.
Recursive self-improvement, in the wild.
Let’s see how far swarm intelligence can take this.
Links below:
🚨 New paper alert !!
🎥 Video VLMs are strong at high-level semantics and long-range temporal understanding.
🧠 JEPA is almost the opposite: better at dense, high-frequency dynamics, local physical consistency, and fast corrective control, but are less suited for rich semantic reasoning and long-horizon reasoning.
We try to get the best of both:
🧩 A VLM as a cortex-like reasoner for semantics and long-horizon planning
⚡ A JEPA branch as a cerebellum-like controller for fine-grained dynamics, physical consistency, and rapid corrections
Proudly, we present ThinkJEPA: a VLM-guided latent world model that FiLM-fuse the pyramid repr of VLMs encoding long-horizon semantic reasoning into the JEPA repr for fine-grained, physically consistent dynamics prediction.
🔗 Project: https://t.co/quro6Pf8un
📄 Paper: https://t.co/yO5rv3ZJT7
@JustinAngel@iclr_conf Thanks so much for the great insight!!
- We are adding SFT / preference based / prompt optimization comparison now🙌 will see what we can offer
- also added qualitative analysis with CBT indicators to compare differences before/after training in the new version
@deltaVee42 Central to that would be building simulation of outcome which is itself as hard as improving therapy skills for chatbot(or even harder than that).
That’s very true. Therapy has a long history and includes many different schools of thought, yet it still doesn’t offer an acute, universally effective solution like antibiotics do for bacterial infections. The human mind is one of the most complex systems we know, and generations of people have devoted their efforts to understanding and supporting it.
What actually changes?
Untrained model:
– warm
– reassuring
– repetitive
– no structure
Trained model:
→ identifies automatic thoughts
→ guides cognitive restructuring
→ does actual CBT
Example:
“What thought flashes through your mind in that exact moment?”
About 1 in 8 young people already use AI chatbots for mental health advice.
But are they actually doing therapy—or just sounding like it?
Introducing 𝐓𝐡𝐞𝐫𝐚𝐩𝐲𝐆𝐲𝐦, a framework to evaluate & align therapy chatbots on clinical fidelity and safety. @eadeli
What does the improvement actually look like?
Untrained: warm, reassuring, repeating, no structure.
Trained: elicits automatic thoughts ("what thought flashes through your mind in that exact moment?"), guide discovery with alternative framing(“What if the fear isn’t being judged, but being not trusted to try again?”).
The untrained model defaults to comfort. The trained model does CBT therapy.
TherapyGym is also a training harness.
CTRS scores + safety penalties → reward signal → GRPO fine-tuning on 13k simulated patient profiles.
Average CTRS: 0.10 → 0.60 (human-rated). Safety violations: 0.38 → 0.20.
But can you trust an LLM to score therapy sessions like a clinician?
We built TherapyJudgeBench: 116 dialogues with 1,270 ratings from licensed CBT practitioners, specifically to audit LLM judges.
Best config reaches Spearman ρ = 0.56 with human raters. For context, the original CTRS inter-rater reliability among human clinicians is 0.59 (Vallis et al., 1986). Our LLM judge is essentially on par.