Are agent skills always worth using?
The answer is no?
This paper provides some important insights to understand this more.
(bookmark it)
Paper summary:
Adding procedural skills to an agent is usually scored by average task success. That number nets gains against damage and hides half of what happened.
Setup: nearly 6,000 paired runs across two office automation benchmarks and three model harness stacks, comparing the same agent with and without skills.
A regression is a task the agent solved without skills and then failed once skills were added. Regressions are large enough that the best performing skills separate themselves mainly through fewer regressions. Larger gains contribute much less.
Three mechanisms drive it.
Skill description osmosis, where a skill changes agent behavior just by sitting in context even when it is never invoked. Grounding displacement, where a prescribed procedure overrides how the agent reads its inputs. Verification displacement, where the procedure suppresses checks the agent would otherwise run on its own output.
Trace analysis surfaces that procedural guidance is the stage least often responsible for failure, while grounding and verification dominate the errors that remain. Existing skills are almost entirely procedure.
Paper: https://t.co/6dNPyN6ali
Learn to build effective AI agents in our academy: https://t.co/1e8RZKs4uX
I don’t understand why FDE is so hyped up!
It’s basically client facing role where you help your client integrate and use AI!
Same shit we used to do at gnani ai back in 2023-2024!
“CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis” has been accepted as a poster at KDD 2026 (@kdd_news) — one of the leading conferences for AI, machine learning, and data mining.
At KDD? Meet with Sentient researcher Darshan Tank (@TankDarshan7) to learn why LLM agents struggle with multi-tool, long-form analysis in crypto research.
Agents can fetch data and nail the facts, but combining it all into real research is where they start to break.
Sentient researchers Darshan Tank (@TankDarshan7) and Sidhant Rahi tackle this in "CryptoAnalystBench: Failures in Multi-Tool Long-Form LLM Analysis", accepted into KDD 2026 (@kdd_news) in Jeju, South Korea ↓
@prathoshap I created skill that help you create poster from your research paper after asking you few questions around your preferences.
Happy to share it via skillpatch
https://t.co/AJq4rOghpe
https://t.co/1ekQh1i424
Everyone assumes more skills = better agents.
But Sentient researchers Darshan Tank (@TankDarshan7) and Baran Nama found that adding skills to an agent can break tasks it used to solve.
In their paper "The Regression Tax", this hidden cost wiped out 59% of the gains that skills produced across nearly 6,000 runs ↓