Excited to share our new work: 🧩 Demystifying Agent Skills 🤖
Agent skills are becoming a key ingredient for LLM agents 🤖🧠 but Why do agent skills actually help—until when they don’t?
Across 8,135 trials, we find the answer is surprisingly simple:
💡 Skills are valuable less as stored knowledge, and more as procedural anchors.
They compress messy experience into reusable ways of setting up, acting, and verifying.
📈 Skills beat workflow memory by +6.06 pts
🧠 65.7% of their value comes from procedural anchoring, while only 4.5% come from supplying missing knowledge 🧠
🔎 As skill libraries grow, retrieval—not generation—becomes the bottleneck
🎯 Retrieving the exact “ground-truth” skill is neither necessary nor sufficient
💥 Skills can still fail when procedures are brittle, context-mismatched, or poorly adapted
✨ Takeaway: Self-improving agents need better abstractions(distilling, retrieving, and adapting the right procedural abstractions), not just more experience. 🧩🚀
📄 Paper: https://t.co/f1KMxUo2yn
🌐 Website: https://t.co/K57RE4ai0g
💻 GitHub: https://t.co/HtaoPPL0MH
Plz upvote if you can !!👇
https://t.co/nMXToqncyE
This work is a collaborative effort with Zhiyuan, @fangruihuang@harvenx01@GaoYipeng@MengdiWang10@Shilong_Liu_AI
#LLM #agent #skills
COLM 2026 Main Accepted🎉!
Is AI good at predicting how groups will behave in the future?
We know AI can make such predictions. But how well can it actually do? And more importantly, how can we make it better?
In our new work Simulating Organized Group Behavior: New Framework, Benchmark, and Analysis we study this problem systematically. Paper link: https://t.co/2k87ulCkx7
If you're interested in AI for future prediction—especially predicting the behavior of companies, organizations, governments, and other organized groups—check out our work:
🔗 https://t.co/zFFXgJkKms
Huge thanks to my amazing collaborators @yeeelow233 , @JasonWuzh@NanHuang99 and other collaborators, and to my advisors @LetianPeng and @shangjingbo for their invaluable guidance
Environment generation is the missing scaling axis for embodied AI.
Introducing SimWorld Studio: a self-evolving factory for endless interactive 3D env where agents act, fail & learn.
Env-agent co-evolvution improves navigation success 50% → 90%.
From a prompt, our SimCoder writes code to automatically build an interactive world. Agents train inside it. And their performance shapes the next world.
On-policy self-distillation is a promising direction for learning from rich textual feedback. But can it really learn from failed trajectories?
Our answer: not quite -- unless we let the model actively interpret them.
🧵1/N
What makes us humans general agents? 🤔
It’s probably not that we’ve mastered every app or UI. New tools appear constantly, yet we become productive quickly. What really transfers are our cognitive abilities: how we perceive, reason, and use memory. 🧠
Introducing CocoaBench, an evaluation framework for general agents with compositional cognitive abilities.
👉 Features
1⃣Complex, realistic tasks🧩
• Human-crafted tasks that are long-horizon, understandable, and span diverse scenarios and domains
• Assume only a small set of general tools (browser, terminal, file system; no per-task APIs)
• Challenging for existing agent systems. ChatGPT Agent reaches only 44% success rate.
• Check out our example tasks on the website (link in reply) and see if you can solve them. They’re fun 🙂
2⃣Covering diverse cognitive abilities 🧠
CocoaBench covers different choices in the following dimensions of cognitive architecture:
• Perception: how the agent gathers and preprocesses information from websites, terminal outputs, files, and images
• Reasoning : planning, deductive/inductive/abductive inference, and both symbolic & visual reasoning skills
• Memory : managing working memory in the long horizon, saving and loading procedural memory (skills), etc.
3⃣CocoaAgent framework🛠️
• Seamless integration with the AIO Sandbox for isolated browser / terminal / file operations
• Easy to plug in any models or your own custom agent
• Flexible evaluation functions
This is just our first release, and we’re actively expanding CocoaBench with more tasks, analyses, and agents. Excited to see what the community builds on top of it!
Project website + more details in the thread👇
🎉 We release DAComp
a full-lifecycle benchmark for data agents across DE-Arch, DE-Impl, DE-Evol, and Data Analysis
combining repo-level coding, open-ended analytics,
with exec-based & LLM-judge evaluation, and visualization
👉 https://t.co/hLplXpNJ9x
👉 https://t.co/6mHYtahVeu
🔍 From simple code completion to autonomous software engineering agents — what changed in the past 5 years?
We wrote the playbook 📖 "𝐅𝐫𝐨𝐦 𝐂𝐨𝐝𝐞 𝐅𝐨𝐮𝐧𝐝𝐚𝐭𝐢𝐨𝐧 𝐌𝐨𝐝𝐞𝐥𝐬 𝐭𝐨 𝐀𝐠𝐞𝐧𝐭𝐬" — 300 pages covering exact recipes 🧪, scaling laws 📈 & RL techniques 🎯 for state-of-the-art Code LLMs.
What's inside:
✨ Full lifecycle: Data → Pre-training → SFT → RL specifically for Code LLMs.
🧪 Empirical Training Recipes: We reveal language-specific scaling laws (Python vs. Java), where Python benefits massively from scale, but C# and Java are "easier" to learn and saturate faster.
🤖 SWE Agents Taxonomy: A detailed look at agents that handle the real tasks, including Environment, Dev, Testing, and Maintenance — moving beyond simple generation to full workflow automation.
This work was led by Jian Yang (@jian_yang96 ) at Beihang University, alongside a stellar collaboration of researchers from Alibaba, ByteDance, and other affiliations. And please find more details in the paper!
🔗 https://t.co/sJwikNRw7H
🚨🚨Can agents earn money, run a business, or even self-organize a society in the physical social world? 🤖🤖
Can agents learn continually to survive and thrive in embodied environments, like how human babies grow? 👶
Super excited to introduce SimWorld, an open-ended simulator of LLM agents in infinite, realistic embodied worlds.
SimWorld features 3 key designs:
1⃣Open-ended realistic world simulation
- built on Unreal Engine 5, with accurate physical social dynamics
- 100+ built-in environments (city, island, wilderness ...)
- language-controllable procedural generation
- text-to-3D asset generation
2⃣Native interface for LLM/VLM agents
- Gym-like agent-environment interaction APIs
- plug in any LLMs/VLMs (GPTs, Gemini, Qwen ...)
- rich multi-modal perception
- open-vocabulary natural-language action outputs
3⃣Diverse physical and social reasoning scenarios
- long-horizon embodied reasoning
- multi-agent collaboration / competition
- easily customizable for any reasoning tasks
SimWorld is fully open-sourced, with a hope to become a foundational infrastructure for real-world agent research across disciplines: robotics, economy, public health, education, etc.
Project website + more details in the thread👇 ...1/
Thrilled to release new paper: “Scaling Latent Reasoning via Looped Language Models.”
TLDR: We scale up loop language models to 2.6 billion parameters, and pretrained on > 7 trillion tokens. The resulting model is on par with SOTA language models of 2 to 3x size.
Can LLMs reason beyond context limits? 🤔
Introducing Knowledge Flow, a training-free method that helped gpt-oss-120b & Qwen3-235B achieve 100% on the AIME-25, no tools.
How? like human deliberation, for LLMs.
📝 Blog: https://t.co/9O7TgH5tJs
💻 Code: https://t.co/hHSLzol4Wg
LiveCodeBench Pro remains one of the most challenging code benchmarks, but its evaluation and verification process is still a black box.
We introduce AutoCode, which democratizes evaluation allowing anyone to locally run verification and perform RL training!
For the first time, we also show that an LLM can act as a problem setter, transforming a simple problem into a harder version sometimes even harder than what it can solve itself.
In other words, LLMs can generate problems they can’t yet solve, opening the door to true self-play.
Moreover, through an agentic framework, we find that LLMs can automatically generate test cases, achieving 98.7% evaluation consistency, which is already highly practical accuracy for an RL verifier.
🚀 Excited to share our work at Bytedance Seed!
Knapsack RL: Unlocking Exploration of LLMs via Budget Allocation 🎒
Exploration in LLM training is crucial but expensive.
Uniform rollout allocation is wasteful:
✅ Easy tasks → always solved → 0 gradient
❌ Hard tasks → always fail → 0 gradient
💡 Our idea: treat exploration as a knapsack problem → allocate rollouts where they matter most.
✨ Results:
🔼 +20–40% more non-zero gradients
🧮 Up to 93 rollouts for hard tasks (w/o extra compute)
📈 +2–4 avg points, +9 peak gains on math benchmarks
💰 ~2× cheaper than uniform allocation
📄 Paper: https://t.co/3HbOwV2tLL
We are glad that TIS and FlashRL have received broad attention from the open-source community that they have been verified and supported (OpenRLHF @hijkzzz, SkyRL @NovaSkyAI, REINFORCE++@hijkzzz, OAT @zzlccc)!
A few updates on our blog and FlashRL package:
(1) more in-depth analysis on TIS (how & why it works) added;
(2) @vLLM 0.10.0 supported!
Feel free to check our updated blog (https://t.co/bp2Mvcj65j, refresh or use chrome incognito mode) and GitHub repo (https://t.co/quU1xLm4SC).
🔥 LLMs can fix bugs, but can they make your code faster? We put them to the test on real-world repositories, and the results are in!
🚀 New paper: "SWE-Perf: Can Language Models Optimize Code Performance on Real-World Repositories?"
Key findings:
1️⃣ We introduce SWE-Perf, the FIRST benchmark for repo-level code performance optimization.
2️⃣ There's a huge capability gap: Expert-level optimization is 4.7x better than the best LLM agent (10.9% vs 2.3% performance gain).
3️⃣ But there's potential! On some tasks, agents can match or even BEAT human experts.
Paper: https://t.co/t2nPzL9Fdh
Home: https://t.co/BHBEFNHrJu
Details in thread 🧵
🚀Exciting to see how recent advancements like OpenAI’s O1/O3 & DeepSeek’s R1 are pushing the boundaries!
Check out our latest survey on Complex Reasoning with LLMs. Analyzed over 300 papers to explore the progress.
Paper: https://t.co/k1HGQTA2kN
Github: https://t.co/VpcNVcEBSg