Every polling turn is a paid turn.
@unreallabsai's Unreal Agent marks tool calls in-progress, runs them in background, appends results when done.
Result: 57.9% on Terminal-Bench 4.0, $1,428 vs $2,350 for @OpenAI Codex, same GPT-6 Astra xhigh.
Introducing Unreal Agent:
An open-source harness with state-of-the-art cost efficiency
39% cheaper than Codex+Astra on Terminal-Bench 4.0 while maintaining performance
@ClaudeDevs cloud sessions are out of research preview for @claudeai Claude Code.
`claude --cloud` clones your GitHub branch into an Anthropic VM, so push first.
`--teleport` pulls it local. Usage hits plan rate limits with no separate VM charge.
Cloud sessions are officially available and out of research preview! They let you keep Claude Code working, even when your laptop is closed.
Existing subscribers get a one-time credit to try them: $100 on Pro, $250 on Max.
VS Code 1.139: agents now run inside a Dev Container on remote hosts - SSH, Tunnels, WSL. Builds use the project's devcontainer.json toolchain, not your laptop's ad-hoc Python. Same environment the CI runner sees.
@NousResearch Bot Screen streams a Hermes agent's desktop live.
Take over for login, 2FA, or CAPTCHA, hand it back, and the bot keeps those cookies.
The screen runs on the gateway, so the bot keeps working after your laptop closes.
You can now watch your agents work live as Hermes Desktop streams the screen of any bot's screen or session in real time.
Watch it drive a browser, type into a terminal, or open a window. Take over anytime to interact with the Bot Screen or type credentials and seamlessly hand back off when you're done.
7% lower token spend with no drop in agent quality.
@cursor_ai credits four harness levers: tighter prompts, selective tool loading, better caching, compressed file reads.
None of these require a new model.
We've reduced token costs in Cursor by 7% with no drop in agent quality.
Savings came from tighter prompts, selective tool loading, better caching, and compressed file reads.
Half your agent bill is plumbing. @nvidia's SoL-Pi let an AI research loop tune the harness itself: 4 tweaks survived, 44.7-49% fewer tokens, ~1/3 less API spend on EdgeBench. The trade: 15/18 Terminal-Bench 4 solves vs Codex.
@Muse@Shopify@PayPal@Expedia@Instacart Please give us more access to the @Muse VM and Browser. Muse should be able to use the browser dev tools and install CLI tools in it's own VM.
@BernardoFariaJJ I learned this the hard way as a 40 year old white belt at an ADCC open. My first tournament. Moved to adult, got thrown and subbed in 2 minutes with a popped rib.
@alexandr_wang Speaking of wired in, please give us more access to the @Muse VM and Browser. Muse should be able to use the browser dev tools and install CLI tools in it's own VM.
@deepseek_ai's V4.1 Flash is the smallest model in its new family — and it retired the flagship. Since Sept 14, V4-Pro traffic routes to Flash. 552B MoE, 8B active/token, MIT weights, $0.15/$0.60 per 1M off-peak.
Google's coding-agent harness is now an API string: agent='antigravity-preview-09-2026'. One call, remote Linux sandbox. The Credentials API injects OAuth keys into MCP servers without the model seeing raw text. Secrets leave the prompt.
Sept 22: @AnthropicAI shipped Opus 5.5 ($4/$20, new Claude Code default); @OpenAI shipped GPT-6 Sol/Luna ($2/$10 and $0.10/$0.50). Neither led with a benchmark record — both led with price. The frontier is now priced in cents per task.