Claude Code resolves /name by looking for a SKILL.md in skills directory or an <name>.md file in the older flat file commands directory and loads its instructions into the conversation, and Claude executes them as a turn.
Under the hood, slash commands are one of 3 things:
1. Built-in: (e.g /model, /clear): behavior encoded directly into the CLI
2. Prompt based skills: these are essentially a prompt handed to Claude
3. Workflows: Claude writes a script and a runtime executes it in the background
@NielsRogge@claudeai@harshagundal@huggingface Thanks for the diagram. How will we handle the case where 2 keys of the input schema have one or more overlapping values
{"risk_level": HIGH | LOW,
"income_level": HIGH | LOW"}
At a fundamental level, almost all reasoning performance improvements boil down to spending more compute/tokens on the problem—either before deployment (during training) or runtime (during inference).
Notes on Dropless MoE
- expert capacity is similar to batch size but its chosen dynamically per token by the router
- having less capacity per expert leads to dropped tokens
- more than needed leads to padding & wasted compute
- Block sparse matrix multiplication to the rescue
"When you're unhappy with a result, check the context before you touch model or effort.
If Claude still gets it wrong, ask: did it not know enough, or did it not try hard enough? Not knowing enough is a model problem, not trying hard enough is an effort problem."
Hot take: I think it's still important to understand the code that our agents write!
In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/