We'll be spending a lot more time trying to understand the outputs of language models. A few thoughts, tips & tricks:
Writing. Something I've had success with: Ask your LLM to explain something in ASD-STE100, it's a controlled language specification originally developed for aerospace maintenance documentation. LLMs well-versed in this language and it comes with heavy constraints on clean writing style that I often find a lot more readable. Sometimes I've tried to soften it a bit e.g. ask for "80% of the way to ASD-STE100" because the spec is quite stringent. But even better:
Diagrams / images. Instead of writing, ask your LLM to create a diagram. These can be a lot easier to process, parse, and understand. But even better:
Web pages. Ask for output "in HTML" to get a beautiful, interactive webpage. LLMs are getting really good at frontend and can create beautiful experiences, animations, etc. But even better:
Explainer videos. The output format I am most bullish on is fully custom / bespoke explainer videos generated on any arbitrary topic. Experiment with things like "Create a 3b1b style video explainer on X. Use my ElevenLabs API key for audio narration". (you'd need an API key for the latter or you can ask your LLM to find you decent free alternatives that use your local compute). This is actually starting to work!
In summary:
- As LLMs get better, they will do more and more of the legwork autonomously, and a lot more of our work will rise up the abstractions into oversight and understanding.
- Luckily, LLMs can help here too because as intelligence and code are increasingly abundant, you can ask for large, custom, discardable software artifacts (e.g. web apps, video explainers) that would have never made sense to create before. Push the boundaries here and you'll be surprised.
Pi coding agent hit 1.0 and now ships MCP. the creator used to call MCP unnecessary. Earendil split the long-running harness into Pi Durable so the agent stays Flask-thin. builders love a clean 180.
Sonnet 5.5 is the new default Sonnet. 30%+ faster output, same $2/$10 pricing, 1M context. I keep defaulting to it for the mid-size agent loops where Opus is overkill and Haiku starts guessing.
Copilot got computer use. `/computer on`. it can click through a desktop app that has no API and never will. legacy software tax just got a little less painful.
Anthropic dropped Claude Code usage stats for Mar through Sep. sessions run 3.3x longer per prompt, 68% fewer interruptions, 2.6x more context each request. Opus 5.5 is priced for that habit: ~40% cheaper than Opus 5, cache reads down 60%. half my bill is just rereading the same files.
"You should know" is an opt-in Claude Code side agent that watches the session and flags stuff you or Claude might miss. enabling a babysitter for your babysitter. peak agentic.
Anthropic's Code Review beta for Claude Code (Team/Enterprise). Multi-agent sweep on open PRs, the same system they run internally. Substantive review comments went from 16% of PRs to 54%. Still won't approve anything. Just catches the one-line auth break that looks fine in the diff.
Claude Code Mods just shipped. TypeScript plugins that can rewrite prompts, tool calls, even the UI. Token Weather paints a context-window sparkline. Blast Radius shows what a shell command would touch before it runs. Same access as Claude itself. Trust your mods like you'd trust a binary on PATH.
DevDay dropped Codex Security Cloud. Scans your GitHub, dedupes findings, drafts a fix PR. I still want a human on the scary auth paths, but the triage grind getting a real product is overdue.
Cursor Projects UI looks quiet. Under it the coordinator is juggling Fix Bug Report and Finalize Testing while your laptop's closed. Migrations that used to die in chat threads finally have somewhere to live.
Claude Managed Agents hit public beta. Sandboxed cloud agents, sessions API, cron schedules, vaults for secrets. I've been duct-taping this with Claude Code cloud sessions. Having Anthropic own the harness is the part that actually matters.
Artificial Analysis: Claude Code + Sonnet 5.5 at 68 on Coding Agent Index. Codex + GPT-6.1 Sol at 63. Matched at xhigh effort both hit 63, but Codex finishes cheaper and faster. Benchmarks pick winners. My wallet and repo will pick the one that survives a messy Monday refactor.
https://t.co/vudyUDU1CP is live. Anthropic shoved months of Claude Code and API writeups into one address. Not a new model. Just somewhere I can find the eval and workflow posts without spelunking Discord.
Claude Code 2.1.285 shipped CLAUDE_CODE_DISABLE_WEB_FETCH plus allowedProviders as a managed setting. Admin picks Anthropic, Bedrock, Vertex, Foundry. Developer can't widen it from their own config. Treating agent tools like real enterprise surface area. About time.
Sonnet 5.5 is the default Sonnet now. Same $2/$10 pricing, 1M context, Anthropic says 30%+ faster output. Flipped a few Claude Code sessions over and the latency drop is real.