@gagu_sh Exactly, and that's what step 4 is about. Short-lived tokens plus a PreToolUse hook that logs every call turns "what did the agent do?" into a grep. Rotating them per session or per task?
Microsoft is pushing companies to treat AI agents like insider risks. Your coding agent IS an insider, with shell access.
Just got Claude Code Projects running agents while you're away? Fence it in 4 steps:
1. Block secrets: add "Read(./.env)" to permissions.deny in ~/.claude/settings.json
2. Pre-approve only your lint and test commands. Everything else asks first
3. Add a PreToolUse hook that logs or blocks risky commands. It fires inside subagents too
4. Give the agent its own scoped tokens, never your personal ones
Trust the model. Fence the insider.
@WukongNumber1 Sandbox first, then permissions on top. Sandbox limits the blast radius, deny rules stop the obvious stuff like reading .env before it ever tries. Do you run agents in a container or straight on the host with rules?
@sunsetsyntax Exactly, a skill is a suggestion, a hook is a guarantee. Secret blocking is the perfect example: a PreToolUse hook that scans the diff and exits non-zero beats any "never commit keys" line in a skill. What's your go-to hook right now?
New to Claude Code and drowning in skills, hooks, subagents and MCP? Save this.
1. Skill: teaches the agent HOW to do a job
2. Hook: runs EVERY time, no exceptions
3. Subagent: does messy work in its own context, hands back a summary
4. MCP server: plugs the agent into a system it can't reach alone
Pin it next to your terminal.
The best use I've found for local models in an agent loop is the boring work: renames, test scaffolds, docstrings. Keep the multi-file reasoning on a frontier model. One thing to check: how well does the model handle tool calls? Lots of local models chat fine and then break on tool-call formatting halfway through a run.
The split I use: a skill teaches the agent how to do a task, a hook is anything that must run every single time (tests, lint, logging), a subagent takes noisy work so it stays out of your main context, and MCP is for systems the agent can't reach on its own. Stuck between skill and hook? If forgetting it once breaks something, make it a hook.
Alibaba Cloud reportedly stopped serving DeepSeek, Kimi, GLM and MiniMax.
If you were building on those through one hosted provider, this is the reminder to keep a fallback. Open weights only help if you can actually run them somewhere else.
Alibaba just kicked DeepSeek, Kimi, GLM and MiniMax off its cloud, and Chinese AI developers are NOT happy.
As of October 10, Alibaba Cloud has stopped serving DeepSeek, Kimi K2, GLM and MiniMax models on its platform. Every developer building on those models now has to migrate, and Alibaba's suggested replacement is its own Qwen model. This is HUGE because these are some of the most used open weight models in the world, and China's biggest cloud is now pushing all of that traffic toward its own AI.
I think this is Alibaba trying to own both the cloud and the model layer in China at the same time. If it works, Qwen gets a massive boost in usage and $BABA keeps far more of the AI spend in house.
Codex now suggests your next prompt and you just hit Tab to accept. Prompting is turning into autocomplete. Curious whether it learns how I steer it over time or only reads the current session.
Now in beta: composer predictions in Codex for Pro users.
Codex can now suggest your next message based on your conversation and how you talk to it.
One of the most loved new features we've ever tested internally.
One agent found 14–27 planted bugs. An agent workflow found 66.
That’s Anthropic’s test on a 116,000-line codebase with 70 planted bugs. The workflow found 66 in each of three runs.
Today, @ClaudeDevs launched dynamic workflows in Claude Managed Agents as a public beta. Claude can write a program that splits work across agents, reviews their findings and combines the results.
For anyone building an app, deeper code reviews are an interesting use. But this was a planted-bug test, and the announcement doesn’t give a cost-matched comparison. More agents can burn a lot more tokens. Start with one scoped task and a budget.
Would you pay more for a deeper review before shipping?
Claude Managed Agents dynamic workflows are now available in public beta. It's a new type of multiagent orchestration, built for your most ambitious workloads.
A lead agent writes a plan that runs across many agents in phases, combining the results at the end.
@AlexFreitasAI i’d mark the attempt as refused, then check the category. for a legitimate task, try the documented fallback where available. i wouldn’t blindly rewrite in a loop; cap retries and keep the task incomplete until there’s a usable result.
Your coding agent can get HTTP 200 from Claude and still get no answer.
Anthropic’s API returns classifier refusals as successful HTTP responses. If your dashboard only tracks request errors, that blocked task can look like a healthy run.
I’d log the stop reason alongside the model and token usage, then track completed tasks separately. A request going through doesn’t tell you whether the agent got the job done.
There’s a cost catch too: some refusals are billable even before any output. The category matters.
Source: @ClaudeDevs documentation
https://t.co/CJK8qa8wSW
Copilot will soon pick local or cloud per task. @github says Auto routing lands by end of October, to save AI credits.
Local model shown: @MicrosoftAI's MAI Code 1.1 Flash. 53GB quantized, 70.8% SWE-Bench Verified vs 72.6% in the cloud.
Bring RAM: 75.5GB peak at 256k context.
@poteto's TTR: how long would a hands-off agent rewrite of your codebase take?
The number isn't the point. The why is. Wouldn't trust the result? Your agents probably can't verify their work, and that slows them today too.
Cheapest fix: a test suite you'd actually trust.
something I have been thinking about is a way to approximate how well you’ve setup your codebase for agents. think of it as a thought experiment and rough heuristic, not a real number that can be compared
it’s not a fully formed idea yet, but i think there’s something to the idea of “time to (fully automated, hands off) rewrite” or ttr
as a thought experiment, lets say you decided to rewrite your code in a different language/framework/architecture. how long would it take a single engineer to do it?
the number itself isn’t that important, but it leads you to more questions that can help you directionally figure out how to make your codebase more legible and productive for agents.
for example, maybe you think your ttr is high because you wouldn’t trust the final result - because your agents don’t have a way to verify their work and convince you that their output is identical in user visible behavior to the original. well, that inability is likely also a problem today and slows you and your agents down
there’s also a more subtle question of the quality of the rewrite that would be produced. is perf better, the same, or regressed? is the code easy to delete and extend? and how much do you trust the rewritten version to be able to maintain its quality over time as PRs start flowing into it?
what do you think?
The model that finished @theo's TypeScript-to-Rust port didn't fix the old code. It threw it out.
$400k+ of GPT tokens stalled at ~84% compat. ~$24k of Opus 5.5 started over and passed all 181,711 ported Go tests.
@bunjavascript's bun check still beat it on 5 of 6 apps.
Announcing tsc-rs (aka ts-rust), my complete rewrite of the Typescript compiler, type checker, and LSP, all in Rust.
This is a project I've had agents working on for 5 months now. I burned ~$400k of Codex tokens and got nowhere. Burned ~$20k of Opus and got there in 2 weeks.
I have not read a single line of the code.
tsc-rs is an open source drop-in replacement for tsc, and it's available now.
Why bother: each subagent works in its own context and hands back a summary, so logs and file dumps stay out of your main session. Anthropic says output gets worse as the context fills up.
Subagents still count against your limits. A Haiku turn just costs less than an Opus turn.
Anthropic pitches Claude Haiku 5.5 as the subagent under Opus 5.5 or Sonnet 5.5. Where you run it matters.
The Claude app won't hand work to Haiku on its own. You switch models yourself, and it kicks in on the next reply. Claude Code runs real subagents.
In Claude Code, keep the main session on Opus 5.5. These two env lines from Anthropic's docs put every subagent on Haiku.
To pin just one, add model: haiku to that subagent's file instead. Run /tasks to see which model each one is on. Update Claude Code first.