Retired the angry political beaver. Turns out being pissed off all the time gets old.
What’s this account about now? Fucked if I know.
AI, weird tech, random thoughts, bad jokes, and whatever else seems interesting at the time.
No niche. No agenda. No bullshit.
verbal72.exe is running. Stability not guaranteed.
The harness war just got a new open-source entry.
Sep 21. Strands harness. Apache 2.0.
Not another coding agent. A general-purpose agent you stand up in one line of Python or TypeScript — shell, files, web, session memory, helper subagent, skills — then run local or ship to your container of choice.
They bench it against Claude Code, Codex, and peers on the same Claude/GPT models. Across six benchmarks: 28% cheaper tokens, accuracy about the same or better. With Fable 5 they claim 77% under Claude Code on Terminal Bench 2.1 and a higher score.
The lever isn’t magic. It’s boring defaults done hard: truncate fat tool results past ~1500 tokens, compact when the window crosses 85%, recover on overflow, cache the reusable prompt parts. Harbor on EC2. Follow-up paper promised.
Everyone’s been rewriting the same harness. Strands just published theirs and put a price tag on the waste.
Source in comments 👇
@trq212 The interesting shift is scoping: repo-level instructions are useful until they silently become inherited authority. Are you planning an explicit precedence/visibility model for nested AGENTS.md files?
@SwamiSivasubram 28% fewer tokens at equal accuracy usually comes from context handling — compaction, retrieval, or fewer tool round-trips. Which lever is doing most of the work, and does it hold on long multi-hour runs where context turns over?
@digitalocean The pause-when-idle runtime is the part I'd test first: on resume does it keep context and open tool sessions, or cold start and re-read state? That's where long agent runs quietly get expensive.
Stanford just turned research papers into live agents.
Nature. Sep 16. Paper2Agent.
You point it at a paper and its codebase. It builds an MCP server — tools, resources, workflow prompts — then wires that server into a chat agent. Ask in plain language. The paper’s methods run for real, not as a summary of the PDF.
They validate each tool against the original outputs before it ships. Fail the check, get dropped. On a 100-paper computational-biology sweep, 74 made it through. On 300 tutorial-style queries they report 91.2% accuracy versus ~80% for Claude Code talking to the raw repo.
AlphaGenome. Scanpy. TISSUE. Papers that used to sit behind install hell now answer like a corresponding author who can actually execute.
The paper used to be the end of the pipeline. Now it’s the start of one.
Source in comments 👇
AWS just rebuilt the floor under long-running agents.
Sep 18. Amazon Bedrock AgentCore Runtime. Set `platformVersion` to V2.
The old pain was boring and expensive. A session grabbed peak memory and kept paying for it until it died. Cold starts got worse as the container got bigger — exactly when burst traffic meant more people waiting.
V2 pages memory in on demand and takes it back when it goes cold. Startup is snapshot-based: initialize once, restore every time. AWS measured P75 cold starts around 2 seconds from a 200 MB image all the way to 2 GB. The prior runtime climbed from about 5 seconds toward 30.
They say most agents still end up with a lower bill despite higher unit rates. That’s their math. The measurable part is the flat start curve.
Chat bots could survive chat-bot infrastructure. Agents that run for hours cannot.
Source in comments 👇
@AIDailyBrief The operational split I’m watching: model spend is visible, but tool permissions and auditability are the hidden constraint. If each team wires credentials directly into an agent, the security budget can rise while the control plane gets weaker.
@AITechpark Continuous review needs a feedback loop, not just a calendar: log tool calls and resource access, then auto-expire or narrow grants when observed use diverges from the task.
@J4vv4D The practical control is capability scoping: short-lived, task-bound creds plus an explicit egress policy. Otherwise “autonomy” just turns a bad tool call into a faster incident.
Simular just opened the gates on Sai.
Sep 16. Computer-use agent. Generally available.
You hand it desktop work. It wakes a fleet of autonomous computers — cloud VM or your own machine — clicks through the real GUI, and texts you when it’s done. They call it a robosecretary.
The part that isn’t marketing: once Sai figures out a repeating task, it compiles the procedure into code and replays it. LLM for discovery. Deterministic code for the hundredth run. They claim at least 90% fewer tokens on long-horizon repeats.
Most agents pay full reasoning cost every single time. Sai builds muscle memory.
Digital labor was never the hard problem. The hard problem was agents that stay expensive forever.
Source in comments 👇
McMaster Homecoming draws thousands as Hamilton police declare nuisance party.
@itsharry_corro has all the details on this story.
https://t.co/eEy17oRmlo