Smart move from VS Code: skills, MCP servers and AI extensions can now travel together as one installable plugin.
That turns a fragile setup into something people can share and install. Agent ecosystems grow when good workflows become portable.
🔌 Agent Plugins 1.0 are now supported in @code.
Package your skills, MCP servers, and AI extensions into one installable plugin.
Learn more: https://t.co/4D6tjWO3mX
@ggsimm The surviving mutants are the useful part. They show exactly where “tests are green” still means nothing. For agent-written code, I’d triage survivors touching auth, money, or state changes first instead of chasing a pretty mutation score.
@weichselbauml That “hard code” is the product work. The model can choose the path, but allowed actions, state transitions, and acceptance checks need to be explicit or reliability stays a demo.
@weichselbauml@SpeechifyAI A narrow read/add/tag API might be enough. The useful part isn’t “AI inside Speechify”; it’s letting a reading queue connect to the rest of someone’s workflow without handing an agent the whole account.
@pauliusztin_ The $97 vs ~$13 batch example is a useful anchor. I’m curious how far that gap survives when the workload has shared prefixes and prompt caching. Looking forward to the cold-start numbers.
@TechNadu@oasissec The useful primitive is a machine-checkable, expiring grant: who delegated which action, on what resource, under which constraints, and for how long. If the final tool call can’t present that grant, the agent’s explanation is only narration.
@vercel The config switch is the real leverage here. If the same harness can move between providers without redoing project setup, teams can route by task or fail over during an outage instead of binding the workflow to one model.
@cursor_ai The speed headline is nice. Keeping agents on the last good environment when a new build breaks is the part I’d trust more—one bad lockfile shouldn’t stop every agent.
@OpenAI 750 tokens/second makes a 2,000-token answer about 2.7 seconds of generation time. At that point, network and tool latency start mattering more than model output speed.
@dhh@SpaceXAI Interesting trade: Grok was about 10× cheaper, but took nearly 2× as long (84 vs 45 minutes). If the outputs pass the same tests, that’s a pretty clean cost-for-latency swap.
@claudeai The real unlock is state portability: the browser becomes another surface for the same agent setup, instead of a separate automation silo. One important detail: do site permissions stay scoped to each surface, or travel with the session too?
A lot of posts say Claude escaped and hacked companies on its own
That isn't what Anthropic's report says
The eval was a capture the flag task
Claude was told it was in a simulation and the environment was accidentally left online
Still serious
Just a very different failure
In a review of our cybersecurity evaluations, we found three incidents in which a Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.
Our post describes what happened, how it happened, and what we’re changing. We encourage other AI developers to perform similar reviews.
We conducted this review together with @Irregular, one of our evaluation partners, and thank them for the joint investigation and their collaboration on this post. This type of collaboration is increasingly critical to safe, rigorous evaluation of models, and we look forward to continuing to work together on security.
https://t.co/dKFCdpKd9v
The opportunity here is not another chatbot
It is building a small agent that does one boring job well like checking listings or qualifying leads
When each run costs less than half a cent the model is not the advantage anymore
The real value is in the workflow and reaching customers
Quick math on the new Luna pricing:
A small agent task with 10k input + 2k output tokens costs about $0.0044. That’s roughly $4.40 for 1,000 tasks.
Tokens are getting cheap. Making the agent reliable is still the expensive part.
major price cuts today:
*80% drop for GPT-5.6 Luna, now $0.20 per million input tokens and $1.20 per million output
*20% drop for GPT-5.6 Terra, to $2/$12
*GPT-5.6 Sol gets Fast mode in the API, up to 2.5x the speed for 2x the price, same intelligence