BREAKING: GLM-5.2 is now 1st on Design Arena.
With an Elo of 1360, GLM-5.2 has jumped ahead of the now unavailable Claude Fable 5.
And it's open weights.
This is an improvement of 4 positions and 27 Elo points to achieve one of the highest Elo scores in our code categories since Design Arena started.
Huge congratulations to the @Zai_org on the release!
Latency is a one-way bridge in user preference.
Slow tokens feel acceptable only until you’ve experienced fast tokens. 3G mobile data felt fine before 4G; after 4G, going back felt broken.
Same with inference. After trying Composer 2.5 in @cursor, I only go back to Opus/GPT for genuinely hard tasks.
For everyday AI work, fast inference is not a nice-to-have. It changes the product category.
Introducing Respan AI Gateway.
The world’s first AI gateway with built-in observability, evals, and prompt optimization.
Access 1,000+ models through one endpoint, with the production layer teams need beyond routing.
Try `ante` by @antigma_labs because it is really good https://t.co/42k1osCHSD I started using it more than @opencode because of just how responsive it is. It just needs MCP and https://t.co/MD7knRTXaA support and it would be perfect
Just discovered today that you can do skill-scoped stop hook in claude code!
It pokes your cc to continue with a skill-specific prompt, e.g. lower resource constraints and rerun a workflow when OOM is detected
Claude Code 2.1.0 is officially out! claude update to get it
We shipped:
- Shift+enter for newlines, w/ zero setup
- Add hooks directly to agents & skills frontmatter
- Skills: forked context, hot reload, custom agent support, invoke with /
- Agents no longer stop when you deny a tool use
- Configure the model to respond in your language (eg. Japanese, Spanish)
- Wildcard support for tool permissions: eg. Bash(*-h*)
- /teleport your session to https://t.co/pEWPQoSq5t
- Overall: 1096 commits
https://t.co/5tP3jJKU2E
If you haven't tried Claude Code yet: https://t.co/4pvrYESSLd
Lmk what you think!
we could build a platform where anyone can contribute agent eval cases, and get paid / recognized if the test case helps improve frontier agent harnesses/models.
New on the Anthropic Engineering Blog: Demystifying evals for AI agents.
The capabilities that make agents useful also make them more difficult to evaluate. Here are evaluation strategies that have worked across real-world deployments.
https://t.co/UD0yGglTU0