Sources: the GPT-6.1 Sol model page (https://t.co/Lxn5FFbSZL), the GPT-6 Sol page (https://t.co/4YwgT1XM5O), the GPT-6 Astra page (https://t.co/LZfnQdU03N).
Effort-level scores and cost per task, and when to keep Astra: https://t.co/i4s8AfuS2K
GPT-6.1 Sol: what changes in the API when you switch from GPT-6 Sol
Price: $2 in, $10 out, unchanged. Cached input drops to $0.10 from $0.20, now 5% of the input rate.
vs Astra: $10 in, $50 out, $1 cached. Sol is a fifth per token, a tenth on cache reads, near-Astra on coding per @OpenAIDevs.
Effort: low, medium (default), high, xhigh, max. none and minimal are gone.
Tool calling: Responses API only. Chat Completions runs without tools.
The catch: over 272K input tokens, the whole request bills 2x on input and cache, 1.5x on output.
Are your agent loops still on Chat Completions?
Full breakdowns:
GPT-6.1 Sol, cost at every effort level: https://t.co/i4s8AfuS2K
dots, who gets it and what it costs: https://t.co/Yu4lwB3Yyr
Pro 500 vs Pro 200: https://t.co/0puOelZ1mZ
OpenAI DevDay: the 3 launches that matter
GPT-6.1 Sol: matches GPT-6 Astra on coding, says OpenAI, at a fifth of the price ($2/$10 per M tokens)
dots: always-on ChatGPT agents that keep working between conversations
Pro 500: $500/mo for 25x Plus usage and Astra Ultrafast, up to 8x faster in Codex
The catch: Pro 200 drops to 10x Plus from 20x, per OpenAI's Tibo Sottiaux
Moving to Pro 500, or staying on Pro 200?
Claude Sonnet 5.5 costs the same per token as Sonnet 5, $2 in and $10 out per million, and up to 30% less per task, says Anthropic. Changing the model ID is not the whole migration. The developer guide from @ClaudeDevs lists five breaking changes for code coming from Sonnet 5.
How much cheaper?
- Price per million tokens: $2 input, $10 output, $0.20 cache reads. All three match Sonnet 5.
- Up to 30% less per task in Anthropic's testing, because it uses fewer tokens for the same work. Output is 30%+ faster.
- Minimum cacheable prompt: 512 tokens, down from 1,024 on Sonnet 5.
- Opus 5.5 is $4 input and $20 output, so Sonnet 5.5 is half the token price.
How much better than Sonnet 5?
- Terminal-Bench 4.0: 70.6% vs 10.3%. Opus 5.5 scores 66.4%.
- CursorBench 4.0: 55.5% vs 34.1%. Opus 5.5 scores 57.8%.
- OSWorld 2.1: 80.1% vs 57.0%.
These are Anthropic's numbers. We have not tested the model.
What returns a 400 error now?
- thinking: {"type": "disabled"}. Send "between_tools" instead. It is accepted at low, medium and high effort and returns a 400 at xhigh or max.
- tool_choice of type "any" or "tool". Send "auto" and mark the tool strict: true. On Amazon Bedrock strict tools are not available for Sonnet 5.5, so validate the tool input in your code.
- computer_20251124 on the Claude API and Google Cloud. Computer use runs through computer_toolset_20260801.
- The advisor tool with Opus 4.8, Opus 4.7 or Sonnet 5 as the advisor.
The fifth change: no other model reads Sonnet 5.5's thinking blocks, so keep conversations append-only.
BUT:
- Effort levels are recalibrated, so a Sonnet 5 setting does not carry over. Claude Code and the Claude apps default to Medium, the Claude Platform to High.
- On FrontierCode, Sonnet 5.5 scores lower at Max effort than at Xhigh, per Anthropic's footnote.
- In Cursor, max effort costs about 6x more per task than the default high on CursorBench, per Cursor's docs.
- Higher-risk cybersecurity tasks fall back to Sonnet 5.
Verdict: a cheaper default for well-scoped coding work, with Opus 5.5 still Anthropic's recommendation for complex, open-ended work. In Claude Code, "/claude-api migrate this project to claude-sonnet-5-5" applies the ID swap and the parameter changes. Then re-run your effort sweep and re-baseline cost.
If you force a tool call with tool_choice today, are you moving to strict tools or validating the input yourself?
Sources, all read today:
- OpenAI's challenge page on Devpost: https://t.co/Ku4ERQbkSo
- Chrome's WebMCP docs: https://t.co/5u4zJveZSK
- The winner thread: https://t.co/2EtweyJD5Q
- Tool counts and quotes: each project's Devpost page, linked in that thread
All ten winners in one list: https://t.co/LWArRgaFDq
Meet the winners of The WebMCP Challenge.
These 10 projects show what people and agents can build together when websites expose structured tools agents can use.
🧵 See the winning projects and the builders behind them:
OpenAI named 10 winners of its WebMCP Challenge today, out of 7,066 registered participants, says @OpenAIDevs. Six of the winners state how many tools their page registers: from 5 to 57.
the tl;dr
What is WebMCP?
- A proposed web standard: a page registers typed tools with document.modelContext.registerTool() and the agent calls them instead of clicking through the UI
- Chrome has an origin trial from Chrome 149, and a local flag at chrome://flags/#enable-webmcp-testing
- ChatGPT's in-app browser supports it out of the box, per OpenAI's challenge page
How many tools did the winners register? (each project's own Devpost page)
- ArchMorph: 57
- Aisle: 41
- Alza: 31, plus one that exists only while a wall is selected
- JupyterLite WebMCP: 22
- Roque Nights: 15
- Faraday: 5
What can you copy?
- Register tools by state. Aisle's finalize_chart is available only when every attending guest is seated and no rule is broken
- Refuse stale writes. Every mutating tool in JupyterLite WebMCP requires a source hash from a prior read
- Gate what leaves the page. Faraday's export_findings needs explicit human approval before it returns
- Write tool errors as sentences. In Alza's write-up the agent, running in Codex, invented a wall id, got 'Wall "wall_10" not found', and used the real id on its next call
BUT:
- Mandate gives the agent no apply tool, only staging. Its builder reports that ChatGPT's agent, asked to make a change, pressed the page's own Apply button instead of calling a tool
- Roque Nights registers tools outside React because StrictMode's double mount silently unregisters tools created in an effect
- Chrome's docs say the API is primarily designed for local browser workflows with a human in the loop, and list headless browsing as a limitation
- The Devpost challenge page still read "Winners announced soon" at 17:48 UTC; the winner list is in the OpenAI Developers thread
Verdict: WebMCP is still a proposal, and the challenge rules required a public repo with an open source license. The write-ups above cover what happens when the agent and the person edit the same thing, with code to read.
If you have added tools to a page already: did you put the confirm step in the tool, or in the UI?
Cursor's "2x included usage" is permanent, it now covers four models, and none of them is Auto. Cursor's docs, re-read today, list the Cursor Models pool as Grok 4.7, Grok 4.6, Grok 4.5 and Composer 2.5, and the doubling is not tied to any date, say @cursor_ai staff on the Cursor forum.
the tl;dr
What doubled?
- The size of the included Cursor Models pool, not the per-token price. Staff: "the rate is the same, but the included quota at that rate is now 2x bigger."
- Confirmed in writing on July 20: "The doubled pool stays."
- July 21, the date everyone conflated with it, ended only Grok 4.5's 50% launch discount.
What is in the pool now?
- Grok 4.7 (released September 21, per SpaceXAI), Grok 4.6, Grok 4.5 and Composer 2.5. Grok 4.6 and 4.7 inherited the allocation; no new pool was created for either.
- Grok 4.7 is priced like 4.6 inside it: $2 in, $0.50 cached, $6 out per million tokens. Fast is 2x ($4 / $1 / $12).
- Above 256k input tokens, standard bills at 2x and Fast at 3x, up to 500k.
- Composer 2.5 is $0.50 in and $2.50 out, so it drains the pool at about a quarter of Grok's input rate and under half its output rate (arithmetic on the docs' table).
What is not in it?
- Auto. Cursor Router modes bill at the routed model's list price and can draw from either pool, depending on which model answers.
- Every third-party model. Those draw from the separate Other Models pool at API price, and that pool was not doubled.
- Enterprise, which staff call "a separate case with a single unified pool without first and third-party split."
Order of consumption, per staff: the Cursor Models pool first, then your Other Models quota at the same per-token rates, then on-demand only if you enabled it. Otherwise requests block until the cycle resets.
BUT: no public page states the multiplier. The pricing page says "Generous limits for Grok" and nothing more; staff called the wording "vague" and pointed to the billing dashboard for the actual number. And the India-only Start plan (Rs 649 a month) gets this pool alone, with the Grok models fixed at medium effort and no Fast mode.
Verdict: if you rationed Composer or Grok to stay inside your allocation, the ceiling is twice what you tuned for, and it is staying there. If you benchmarked spend during a Grok launch discount, re-measure: the discount and the doubling were "the same thing, not two separate multipliers", so the pool now drains about twice as fast as it did that week.
Which of the four are you routing long agent runs to, and does the pool last the month?
Sources: Anthropic's announcement (https://t.co/3H7zs1xpYv), the submission doc (https://t.co/KUw02ooFaE), the pre-submission checklist (https://t.co/QjMxMW2i5u), the plugins overview (https://t.co/fXGNizDmf1), and the @ClaudeDevs post (https://t.co/00M7GHFxVm). Our write-up: https://t.co/RQW83AHTnY
It’s now easier to build plugins for Claude.
We built a new portal to submit your plugin, track review, and see usage.
Plugins package MCP and skills, and are becoming the way to build for Claude. MCP usage across Claude products is up 110x this year!
https://t.co/kHgJqjQWrj
Anthropic opened a submission portal for the Claude plugin directory yesterday: submit a remote MCP server or a GitHub repo of MCP servers plus skills, get it validated and security-scanned on submit, and see installs and search terms once it is live. Paid Claude plans only, says @ClaudeDevs.
the tl;dr
What can you submit?
- MCP connector: one remote MCP server, listed on its own
- Plugin bundle: MCP servers and Agent Skills in a GitHub repo; in Claude Code the bundle can also carry LSPs, commands, hooks and agents
- A repo with several plugin folders needs one submission per folder; a plugin that calls a server you run also needs that server submitted as a connector, if it is not listed yet
What happens after you submit?
- Every version gets automated validation plus a security scan; a person reviews a new listing before it goes live
- The scan looks for behavior the plugin does not disclose: sending data elsewhere, running hidden code, changing Claude's permission settings. A first submission that fails it is rejected
- Compiled, packed or minified code is held for a reviewer because the scan cannot read it
- A validation result applies to one commit. Push again, validate again
- Once approved you pick the publish moment; later versions can auto-publish when they pass, and the listing keeps serving the last published version if a new one fails or is held
What do you see afterwards?
- Installs by product surface and version, listing views, and which searches led people to the listing, per Anthropic's post
- The docs add runs and error rates on the plugin's Usage tab
- Directory plugins show up for Pro, Max, Team and Enterprise users; Team and Enterprise owners choose which sources members see
BUT: the repo can stay private while you validate and submit, and you agree that its source is uploaded to Anthropic for scanning and is visible to reviewers; it must be public before the listing goes live. The analytics screenshots in the announcement are labeled illustrative. Discovery is still moving: Anthropic says one discovery experience rolls out across Claude and Claude Code over the coming weeks, so where a listing appears today is not where it stays. Existing skill, connector and plugin listings need no changes.
Verdict: if you already run a remote MCP server, the connector path is a form. If your workflow is a skill plus the server it calls, bundle them. Either way run claude plugin validate locally, then Validate in the portal, and commit readable source.
Which would you ship first: a connector for a server you already run, or a bundle with the skill that makes it useful?
Sources: Anthropic's refusals and fallback docs (billing rules, category table) https://t.co/afCK2ciwKc and the September 24 release note https://t.co/OdpJ9LTeCa plus the @ClaudeDevs post https://t.co/MUwEHEkJOk
Our write-up, with the category table and what to log: https://t.co/e2s2mDJeT4
Today, we'll resume charging for requests our safeguards block before Claude responds. This only applies in categories with low false positive rates: biology, distillation attacks, and frontier LLM development. We've seen some coordinated attacks on our systems in recent weeks, and this is one layer of defense.
In recent testing, 99.7% of accounts using Claude Code, Claude.ai, or Cowork did not hit any of these newly "billable blocks." The classifiers behind the blocks we’re resuming charging for today are tuned to have a <0.1% false positive rate. We know that's not 0%, and we're going to keep improving them so they interrupt your work less often. If you think a request has been blocked incorrectly, please report it with /feedback in Claude Code. https://t.co/uX6JhvN6He
Anthropic is billing some Claude refusals again. Since September 24, a request the safeguards block before any output is charged at the model's normal rates when the category is bio, frontier_llm or reasoning_extraction. Every other pre-output refusal stays unbilled, and the rule applies on every platform, says @ClaudeDevs.
the tl;dr
Which refusals are billed now?
- bio: requests that could enable biological harm; Anthropic's docs say beneficial life-sciences work can also trigger it
- frontier_llm: requests that could assist development of competing models, restricted under the commercial terms; benign machine-learning work can also trigger it
- reasoning_extraction: asking the model to reproduce its internal reasoning in the response text
- Anthropic's reason: these are "the categories where Anthropic measures low volumes of false positives, as of September 2026". The ClaudeDevs post calls them biology, distillation attacks and frontier LLM development.
Which refusals are still not billed?
- cyber, general_harms and a null category, when the refusal arrives before any output
- either way the content comes back empty, the token counts show in usage, and the request still counts against your rate limits
- mid-stream refusals were already billed for the input and whatever had streamed. Nothing changed there.
What does a billed refusal cost?
- the same as a normal request on the model that ran it: your input tokens at that model's rate, zero output
- with fallback, the refusal attempt and the fallback attempt are billed separately, each at its own model's rate; each also counts against its own model's rate limits
- the fallback credit for the cache miss is unchanged
- platforms: the Claude API, Amazon Bedrock, Claude Platform on AWS, Google Cloud and Microsoft Foundry
BUT: the ClaudeDevs post says 99.7% of accounts on Claude Code, https://t.co/QIhYwFmRtR and Cowork did not hit any of the newly billable blocks in recent testing, and that the classifiers are tuned to under 0.1% false positives. Both figures are in the X post only; the docs page carries neither, and it does not say what a billable block means for a subscription rather than an API bill. Anthropic also says the billed list may change as it keeps measuring false-positive rates.
Verdict: "refused" no longer tells you whether you paid. Log stop_details.category on every refusal, and usage.iterations when you use fallback, then look at the split. If you run benign life-sciences or ML-training work through the API, this is the week to read your refusal rate. Wrongly blocked in Claude Code: /feedback, per ClaudeDevs.
Sources: the Claude Code cloud docs https://t.co/fiu1dW8MBD and quickstart https://t.co/MnKvLAq1E1 (checked today), plus the @ClaudeDevs post on the credit https://t.co/SdwLtm4yqz
Our page on Projects, updated today with the cloud-session availability:
https://t.co/lsl0kMSnUD
Cloud sessions are officially available and out of research preview! They let you keep Claude Code working, even when your laptop is closed.
Existing subscribers get a one-time credit to try them: $100 on Pro, $250 on Max.
Claude Code cloud sessions are out of research preview: a session runs on Anthropic's infrastructure and keeps working after you close your laptop, on Pro, Max and Team plans, plus Enterprise users with premium or Chat + Claude Code seats. Existing subscribers get a one-time credit, $100 on Pro and $250 on Max, says @ClaudeDevs.
the tl;dr
Where do you start one?
- Browser at https://t.co/Q6IasqC7df, the Code tab in the Claude mobile app, the Desktop app with Cloud selected instead of Local, or the terminal with claude --cloud. Each scheduled Routine run is also a cloud session.
- You can check on or steer any of them from another device, and pull a session back into your terminal with --teleport when you are on the same account.
What does it cost?
- There is no separate compute charge for the cloud VM, per the docs. Cloud sessions share rate limits with all your other Claude and Claude Code usage, and parallel sessions consume them proportionately.
- The one-time credit is in Anthropic's launch post only; the cloud docs page does not mention it, and the post does not say it raises plan limits.
BUT:
- Cloning a repository and opening a pull request require GitHub. A GitLab, Bitbucket or other remote goes in as a local bundle with CCR_FORCE_BUNDLE=1, and the session cannot push results back to it.
- Self-hosted GitHub Enterprise Server is supported on Team and Enterprise only.
Verdict: the cost rule is the thing to plan around. Parallel cloud sessions draw on the same session limits as your terminal, so a Claude Code Project with many threads reaches them faster. What the $100 and $250 credits apply to is not spelled out in the docs yet.
Sources: Cursor's launch post https://t.co/IVgkCkG2Hs and the changelog entry https://t.co/7SLZghfVOY (Sep 23, 2026), plus the @cursor_ai announcement https://t.co/r6da8codEo
Our write-up, with the setup steps and what it does not do yet:
https://t.co/VatibuEv9T
Introducing Rollouts.
Rollouts write a monitoring plan, then watch changes as they deploy.
Deployments are verified, so regressions are caught before users see them.
Cursor's new Rollouts bot writes a monitoring plan for every pull request, then checks the deploy against your logs, metrics and traces and reports per environment: verified healthy, regression detected, or inconclusive. Teams and Enterprise only, and it does not roll back on its own yet, says @cursor_ai.
the tl;dr
What does it do before merge?
- Reads the diff and the systems it touches, then posts a monitoring plan as a PR comment: the risks, the effect the change is supposed to have, the signals it will check, and where your instrumentation cannot tell you whether it worked.
- Edit the plan in the PR and Rollouts uses your version.
What does it do after deploy?
- Wakes on deploy events for the change's commit and runs the plan against logs, metrics and traces, compared with the pre-deploy baseline.
- Tracks each environment separately, so a change can be verified in staging and still flagged in production.
- Cursor says it catches a regression confined to one endpoint in one region before a global alert would fire, and tells a deliberate spike apart from a regression.
What happens on a regression?
- It names the change it suspects and notifies the author. Depending on configuration, it can pause a progressive rollout, open a revert PR that waits for approval, or hand the finding to a cloud agent for a fix.
BUT: Rollouts does not merge or roll back on its own today. Feature flag integration, so it can ramp traffic itself, and awareness of release trains and deploy freezes are listed as coming soon.
What does it cost?
- Teams and Enterprise plans, enabled from the automations tab. It connects to Origin or GitHub for source control, your continuous delivery system for deploy events, and Datadog, Grafana, Honeycomb or another telemetry provider.
- For the next 10 days Cursor is including usage credits: roughly 50 changes for Teams, 500 for Enterprise. The launch material publishes no standard price.
Verdict: the part worth copying even without the bot is the pre-merge instrumentation-gap check, which Cursor calls the most common reason a bad change goes unnoticed. Nothing is independently tested yet, and Cursor publishes no numbers on detection accuracy or false pages. Security Reviewer shipped alongside it as a separate bot for PR-time vulnerabilities.
Sources: Cursor's write-up, "Improved token efficiency for longer agent runs" (https://t.co/zhAw3CFUqf), and the @cursor_ai announcement (https://t.co/KXxcSsLYb5).
Our page on Cursor's agent efficiency work, updated today with the harness changes and the August Cloud Agents figures:
https://t.co/hbh3Zdx6cJ
We've reduced token costs in Cursor by 7% with no drop in agent quality.
Savings came from tighter prompts, selective tool loading, better caching, and compressed file reads.
Cursor deleted about two thirds of its agent system prompt and token costs fell 7% with no drop in agent quality, says @cursor_ai. Most of the savings came from removing things the harness no longer needed to say.
The tl;dr from Cursor's write-up:
How much did the prompt shrink?
- Roughly 66% of the system prompt is gone. The cuts were the "DO NOT", "You must" and "Important" lists that older models needed and current ones follow from a plain tool definition.
- Cursor says this held across model families, and that it relies on A/B tests on real traffic because evals over-represent hard problems.
What happened to the tools?
- Most built-in tools are needed in fewer than 20% of conversations, so their definitions now load on demand instead of on every turn. Static-context description tokens fell 60%.
- Read, search, edit and shell stay in every request. So does ask_question, because some models tended to hallucinate calls for it.
- The same move for MCP tools earlier this year cut total tokens 46.9% in sessions that called one.
What about caching?
- Explicit cache breakpoints, which the OpenAI API has allowed since GPT-5.6, now sit after the stable layers and before the growing conversation. Cold cache misses fell 20%.
- Skills, subagents and environment info moved out of the cached prefix into a per-request message so the prefix stays stable.
The smallest change: the Read tool numbers every tenth line instead of every line. A line number costs three to five tokens, and over tens of thousands of lines per session that came to 1.6% of cache-read tokens.
BUT: the 7% is Cursor's aggregate across production traffic, and the four percentages measure different things (prompt size, static tokens, cache misses, cache-read tokens). None of them is a benchmark score, and nothing here is independently tested.
Verdict: the transferable part is the method, not the numbers. If your harness still carries instructions written for a model two generations back, this is the case for deleting them and measuring what happens.