Excited to share our work on understanding when and why agent skills work and when they donโt. Across 8,135 trials, we uncover what makes skills effective, where they break down, and why retrieving and adapting the right procedure matters more than accumulating experience.
Excited to share our new work: ๐งฉ Demystifying Agent Skills ๐ค
Agent skills are becoming a key ingredient for LLM agents ๐ค๐ง but Why do agent skills actually helpโuntil when they donโt?
Across 8,135 trials, we find the answer is surprisingly simple:
๐ก Skills are valuable less as stored knowledge, and more as procedural anchors.
They compress messy experience into reusable ways of setting up, acting, and verifying.
๐ Skills beat workflow memory by +6.06 pts
๐ง 65.7% of their value comes from procedural anchoring, while only 4.5% come from supplying missing knowledge ๐ง
๐ As skill libraries grow, retrievalโnot generationโbecomes the bottleneck
๐ฏ Retrieving the exact โground-truthโ skill is neither necessary nor sufficient
๐ฅ Skills can still fail when procedures are brittle, context-mismatched, or poorly adapted
โจ Takeaway: Self-improving agents need better abstractions(distilling, retrieving, and adapting the right procedural abstractions), not just more experience. ๐งฉ๐
๐ Paper: https://t.co/f1KMxUo2yn
๐ Website: https://t.co/K57RE4ai0g
๐ป GitHub: https://t.co/HtaoPPL0MH
Plz upvote if you can !!๐
https://t.co/nMXToqncyE
This work is a collaborative effort with Zhiyuan, @fangruihuang@harvenx01@GaoYipeng@MengdiWang10@Shilong_Liu_AI
#LLM #agent #skills
@LChoshen And you did focus somewhere different, which is fair, the implications are the interesting part. The one step I'd resist is "improvements don't generalize", one coding benchmark can't carry that. Glad to have the chance to make that clear. Thanks for sharing it again!
@LChoshen Both, honestly, and the split is clean. "New models improve more than expected on hard tasks" holds, that's the +0.40 that survives our controls. "Easier tasks improve less than expected" is the ceiling doing the work: easy items sit at 0.83, so there isn't much room left.
@LChoshen@KumailAlhamoud@xiang_lorraine@YuexingHao Worth deciding early in any collection design, since retrofitting is the expensive part: anything fixed before the models were run. Even on a subset. That's what makes a level shift separable from a shape change later. Dm u more details, happy to chat!!
@LChoshen@KumailAlhamoud@xiang_lorraine@YuexingHao On EEE side: per-example is necessary, aggregates alone can't identify item difficulty. But the binding constraint is where difficulty comes from. Estimated from the same responses you're measuring, IRT goes circular. Human anchors break that, and they're expensive to retrofit.
@LChoshen@KumailAlhamoud@xiang_lorraine@YuexingHao Personally, and it's definitely the thing I want to keep working on. The author list spans several institutions, it came together around the question rather than out of one lab, which suits this problem. Would love to take it further with you.
@LChoshen@KumailAlhamoud@xiang_lorraine@YuexingHao On what's next: the agentic case is still unidentified, since model era and scaffold era move together. Breaking it needs dated models run against dated harnesses. Data and code: https://t.co/ghVn7eTwsK
๐๐๐ Give AI agents personas!
New research on world simulation from Harvard, MIT, and Stanford. We introduce ๐ ๐ฎ๐๐ฟ๐๐๐ , a population-scale simulated-user evaluation infra for testing AI systems and digital products.
๐ ๐ฎ๐๐ฟ๐๐๐ : ๐ฆ๐ถ๐บ๐๐น๐ฎ๐๐ถ๐ป๐ด ๐๐ต๐ฒ ๐ช๐ผ๐ฟ๐น๐ฑ ๐๐ถ๐๐ต ๐ด.๐ฏ ๐๐ถ๐น๐น๐ถ๐ผ๐ป ๐ฃ๐ฒ๐ฟ๐๐ผ๐ป๐ฎ ๐๐ด๐ฒ๐ป๐๐ ๐
๐ Project and Community: https://t.co/trvAPK5H0J
๐ Paper: https://t.co/MktjtmriJJ
๐ฌ Demo: https://t.co/hadWrE9QlH
๐ป GitHub Repo: https://t.co/50AUdIHX3O
๐ฅ Persona 8B: We created 8.3B persona records at the scale of the global population, across 1,290 dimensions spanning background, psychology, capabilities, behavior, and lifestyle.
๐งช MatrAIx Playground: Supports persona-agent evaluations across four environment types: Survey, AI Chatbot, Web, and App.
๐ MatrAIx Applications: Offers 1,000+ evaluation tasks across 25+ domains, including Commerce, Software, Finance, and Healthcare.
๐ค Join our open-source research community (https://t.co/dzcLWYwqO2) of 200+ contributors, including 40+ from OpenAI, Anthropic and Google DeepMind!
@XiaominLi98@YuexingHao
#AI #LLM #Agents #Persona #Evaluation #Infra #UserSimulation #Harvard #MIT #Stanford #OpenAI #Anthropic #GoogleDeepMind #xAI
RAG gives agents knowledge. Tools give agents actions. Fine-tuning adapts their weights.
But none of these give agents *procedural knowledge* โ how to reliably execute specialized tasks step by step.
We propose Agent Skills as the missing layer
LINK: https://t.co/BnlH82Y9jk
๐จ New paper alert !!
๐ฅ Video VLMs are strong at high-level semantics and long-range temporal understanding.
๐ง JEPA is almost the opposite: better at dense, high-frequency dynamics, local physical consistency, and fast corrective control, but are less suited for rich semantic reasoning and long-horizon reasoning.
We try to get the best of both:
๐งฉ A VLM as a cortex-like reasoner for semantics and long-horizon planning
โก A JEPA branch as a cerebellum-like controller for fine-grained dynamics, physical consistency, and rapid corrections
Proudly, we present ThinkJEPA: a VLM-guided latent world model that FiLM-fuse the pyramid repr of VLMs encoding long-horizon semantic reasoning into the JEPA repr for fine-grained, physically consistent dynamics prediction.
๐ Project: https://t.co/quro6Pf8un
๐ Paper: https://t.co/yO5rv3ZJT7
SkillOrchestra (https://t.co/iZQvlzAFNb) hits on something I've been thinking.
The move: learn skills from execution, then match skill demands to agent competence. 700x cheaper training, 22.5% better performance than SOTA RL methods.
Sometimes subtraction beats addition.
@barry_zyj Really like this framing๐. Weโve been writing a paper around a very similar question, whether Agent Skills actually address key bottlenecks of current LLM agents, especially around context efficiency and procedural knowledge.
๐ SGLang GTC Giveaway โ 20 FREE Passes!
SGLang is an open-source LLM serving engine that helps models like DeepSeek, Qwen, Kimi, Minimax, GLM, and Llama run efficiently at production scale.
Thanks to our sponsor @radixark, we're giving away 20 NVIDIA GTC 4-day exhibit passes (worth $930 each)! ๐๏ธ
To enter the lottery:
1๏ธโฃ Follow us โ @lmsysorg
2๏ธโฃ โญ Star SGLang on GitHub โ https://t.co/UKowWNK0Ic
3๏ธโฃ Reply with: your favorite open-source model and what you use it for
4๏ธโฃ Repost this for extra visibility
How we pick winners:
๐ Top 5 most engaging comments win directly
๐ฒ Remaining 15 drawn randomly via xpickr
We'll verify your GitHub star before sending tickets, so make sure you've starred the repo!
Let's go ๐