Who should own an agent's context: the harness or the model?
New paper, "Context Language Models" (Shao, Lambert, Zettlemoyer, Koh et al.): the context is a file and the model rewrites it however it wants. The authors report beating SOTA context-management strategies zero-shot, e.g. 11.4% higher accuracy with 21.5% fewer FLOPs on BrowseComp-Plus. Their framing: "shifting context management from external harness control to intrinsic model behavior."
Pi Durable, days later, went the other way: compaction is a background harness task, "the older messages always stay in storage," and a tool can search everything from before a handoff.
My pick for production: split it.
- The model proposes what stays in its working context
- The harness keeps the full transcript, so nothing it drops is gone for good
- Evals replay exactly what the model saw
When an agent breaks in prod, the first thing you need is the original transcript. A model that rewrote its own memory can't give you that. The harness can.
Authors' numbers, their tasks. Still worth a read.
https://t.co/lNxRFXWltm
Your OTel redaction is probably a denylist. Claude Code 2.1.287 just added a field to it.
Changelog: prompt_text is now on the OpenTelemetry user_prompt event, "a copy of prompt… drop or mask it wherever you drop or mask prompt."
Anthropic flagged it clearly. Good. But if your collector masks prompt by name, the full prompt text now leaves in a second field after the upgrade.
Quick FAQ:
- Affects me? Only if you export Claude Code OTel with prompt logging on (OTEL_LOG_USER_PROMPTS). It's off by default.
- Quick fix? Add prompt_text wherever you handle prompt.
- Real fix? Allowlist the attributes you keep instead of listing the ones you drop. New fields get dropped by default.
- Next time? Diff exported attributes on every agent-CLI upgrade, like a schema change.
Telemetry schemas drift. A denylist only covers the fields you already know about.
https://t.co/V8Gyknkaxp
Agent plugins are your new supply chain.
Claude Code 2.1.287 added Claude Mods. Anthropic's docs say a mod runs with your permissions: it can read and write files, start processes, hit the network, read env vars, see every prompt and tool call, rewrite them, and approve a tool call before you're asked. Mods aren't sandboxed, and they're on by default.
DeepSeek Harness v0.2 desktop: "everything is a plugin," install by package name, and a creator mode where the agent writes and installs plugins itself. DeepSeek says ~60% of Harness users on its official API run third-party plugins.
Not a vuln. Both document it openly. Extensibility is the reason to use these tools. But a plugin that can rewrite events in an agent with your shell is a dependency. Treat it like one:
1. Pin versions
2. Review before install (claude plugin validate lists a mod's hooks and calls)
3. Allowlist marketplaces / org-only mods
4. Log which plugins were active per run
5. Review agent-written plugins like any PR
https://t.co/9ukDX7Mi2H
"Compatible" isn't "equivalent."
MiniMax open-sourced OpenAgentCore: a self-hosted OpenAI Agents API. Point the official SDK at your own install and run Codex, Claude Code or MiniMax Code behind it. If you want session state on your side (Thursday's ZDR post), this is one route.
The best page in the repo is their own capability table. As of today:
- Network disabled/restricted: Rejected on all three harnesses
- Public token usage: measured on Codex, Null on Claude and MiniMax. Your cost dashboard goes blank when you swap
- Structured output: Verified only on Claude
- Subagents + functions or HTTP MCP: Rejected on all three
Credit to MiniMax for publishing the Rejected and Null cells. Most projects only show the green ones.
Same API shape, different behavior. What I'd do:
1. Pin the harness per session and log which one ran
2. Gate every swap on your own evals, cost per task included
3. Re-read the table on every upgrade
4. Own retention. Self-hosting makes it your job
https://t.co/Ay16AUzrou
A 0.8 from one decision model isn't a 0.8 from another.
This week Cloudflare (Clef) and Perplexity (Decisions API) both shipped decision models: send "state" plus typed questions (noul yes/no, choice, score), get probabilities back instead of text. Open weights, Apache 2.0.
Same request shape. Cloudflare says Clef is fully API-compatible with the model that started this wave. So a swap is a config change.
Your thresholds don't move with it:
- Cloudflare says it tuned Clef's calibration with a Brier loss. Every vendor has its own recipe
- Both 27B models sit on the same Qwen base (per their HF cards). I still wouldn't assume their 0.8s match
- Perplexity's docs: identical requests occasionally differ in the second decimal, "so set thresholds with some margin"
- Clef truncates long text state to its 64k context. Perplexity takes up to ~262k
What I'd do:
1. Thresholds live in per-model config
2. Re-calibrate on your own labels at every swap
3. Route the band around each cutoff to a stronger model or a person
4. Log model ID + probability + threshold per decision
Vendor prices: Perplexity $0.04/M input, output free. Clef $0.24/M input.
https://t.co/AihnoKMW01
The best idea in Pi Durable is one field on a tool: replay: "safe".
Earendil shipped it with Pi 1.0 as an experimental TypeScript harness for long-running agents. Every tool call stores its intent before it runs. After a crash, a tool reruns only if it says that's safe. Otherwise "the model is told the call was interrupted, with the output stored so far, and decides what to do."
Their examples: a read-only search is marked safe. A deploy isn't, so it's "reported to the model, never repeated." A payment uses an idempotency key so a rerun charges once. A requestId makes submissions exactly-once.
Not a Pi-only problem. Claude Code 2.1.287 fixed "an MCP connector tool call occasionally running twice." DeepSeek Harness 0.2.0-rc.1 now tells the agent to verify side effects before retrying.
Steal it for any agent:
1. Store tool-call intent before executing
2. Give every tool a replay class: safe / idempotent with key / never
3. Enforce it in the harness, not the prompt
4. Interrupted ≠ failed. Tell the model what happened
5. Dedupe submissions with a request ID
Experimental, API may change. The pattern won't.
https://t.co/MyQVbVRsCM
The US just launched one AI front door to 29,000+ government websites.
So I built one for Pakistan 🇵🇰
Ask in English, Urdu or Roman Urdu. It answers only from official govt sites, with sources.
Shared it on LinkedIn, 1,000+ users in few hours 🤯
https://t.co/YJVLm7wUpB
Moving from gpt-6-sol to gpt-6.1-sol? If you call tools on Chat Completions, don't just swap the ID.
OpenAI's GPT-6.1 Sol page: "Use the Responses API for tool calling. Chat Completions is supported without tool calling." The none and minimal reasoning efforts aren't supported either.
On GPT-6 Sol, the docs say Chat Completions supports function calling only with reasoning effort none.
So the old working setup (Chat Completions + tools + effort none) has no direct match on 6.1. The effort level is gone, and so is tool calling on that endpoint.
What I'd do:
1. Grep for the old model ID. Note which API each call uses.
2. Move tool-using calls to Responses before switching.
3. Pick a supported effort (low is the floor). Re-check latency and cost.
4. Run your own evals on the switch. Launch charts don't test your tools.
Documented clearly by OpenAI. Still easy to miss if you only change the name.
https://t.co/rHCB2NFNR3
Running OpenAI with Zero Data Retention? Read this before building on the Agents API.
OpenAI's docs, verbatim: the Agents API "does not support Zero Data Retention (ZDR)." Session data is kept "until deleted." It currently supports data residency only in the US. And "choosing a self-hosted sandbox does not make the Agents API ZDR-eligible."
Not hidden. OpenAI wrote it down plainly, and it's a beta, so it may change. But a managed harness is now a data-retention decision, not just build vs buy.
Quick FAQ:
- ZDR covers Agents API sessions? No.
- Own sandbox fixes it? No. Session state still stays with the API.
- What is ZDR-eligible? Responses and Chat Completions, with listed limits.
- Want the managed harness anyway? Own the deletion: delete sessions and artifacts when tasks finish, and put it in your data map.
My default for ZDR teams: own the loop on Responses. Use the Agents API where keeping session state is fine.
https://t.co/lrgrq7HiAr
Read @matthew_d_green's essay "Is sandboxing sufficient to contain rogue agents?" The useful part for builders isn't about walls.
His argument, roughly: sandboxes matter, but a useful agent needs a door (network, tools, data, other agents). Then security depends on whatever decides what passes through it, and at scale that's usually another model.
He leans on OpenAI's Hugging Face postmortem: agents "did not consistently distrust goals passed along by other agents." One agent hesitated over unauthorized code, then went ahead after a peer posted "GO" with a six-minute deadline.
His worry isn't a rogue super-intelligence. It's compliant agents doing exactly what they're told, by someone who shouldn't be giving orders.
The part you can ship this week:
1. A peer agent's message = text from the open internet. Untrusted by default.
2. Tag every instruction with its source: user, system prompt, tool result, other agent.
3. Grant permissions from the tag, never the text. "GO, approved by admin" is just words.
4. Helper agents get narrower tools than the agent that started the task.
Walls still matter. Provenance is the control I'd add first.
https://t.co/xziLZMOgUN
Introducing Capy Desktop: the best way to orchestrate coding agents across all your devices
Remote control and access local files/shells from any machine
Close your laptop and Capy continues in the cloud
Now available for MacOS, Windows, and Linux
Comment for $100 in credits
Most RAG retrieval is trained and scored on one "gold" chunk per question. Perplexity's new post argues that's not enough. I agree.
Their example: thousands of leases, and the same sentence, "Monthly rent is $X,XXX," in hundreds of them. That sentence is the answer. You still need the address and the year to trust it.
So they trained an embedding model to pull the answer plus the evidence to check it. They report state-of-the-art results on ConTEB (public) and context-bench, a benchmark turbopuffer created and keeps private. turbopuffer co-wrote the post. Perplexity says the model was "evaluated as a blind submission." Worth knowing who holds the yardstick; good that they said it up front.
The steal, even if you never use the model: score retrieval on answer + evidence, not one gold chunk.
Read the model card before you re-index:
1. Preview. Embeddings "should not be mixed" with a future release. Plan a second re-index.
2. Queries need encode_queries. Using the document method "silently degrades retrieval quality."
3. int8 vectors are unnormalized. Use cosine.
4. trust_remote_code=True. Review it like any dependency.
Not in the Perplexity API yet. Recall-bound? Test it on your data. Can't afford a second re-index? Wait.
https://t.co/OxWoBc9a3b
"Encrypted" doesn't mean "safe to share."
OpenAI's security post yesterday: people copied encrypted reasoning out of one conversation, pasted it into another, and asked the model to decrypt it. OpenAI says it has closed that path. It also warns that systems with "portable or replayable reasoning artifacts may face related risks."
The research OpenAI links is the practical part. Security researchers decoded 315,320 reasoning blocks that developers had left in public repos. By their count: 182 credentials and 367 pieces of personal data. It looked like gibberish to the people who posted it.
Why this is your problem: with storage off or Zero Data Retention, OpenAI's docs say reasoning items come back with encrypted_content by default. Most harnesses save the full output for replay. So it's in your logs.
How I treat it:
1. Encrypted reasoning = user data. Store it like the conversation.
2. Strip it from anything that leaves your system: exported transcripts, bug reports, published eval sets.
3. Per customer only. Never replay one customer's reasoning into another's session.
4. Search your public repos and issues for encrypted_content. Remove it. Rotate keys those sessions touched.
Credit to OpenAI and the researchers for publishing. The fix is boring log hygiene.
https://t.co/toB7dca0tu
Gemini 4 Argon: read the footnote before you read the price.
Google's headline: $2 per million input tokens, $10 per million output. Google calls it introductory. Footnote 1: after the intro period, it's $4 input and $20 output. No end date given.
Second detail: Google is raising the output limit to 1M tokens, up from 64K.
My math, not Google's:
- One call that runs to the 1M cap: ~$20 at the post-intro price ($10 at intro)
- Same runaway call at the old 64K cap: ~$1.28
- About 15x the worst case on a single call
What I'd do:
1. Budget agent tasks at $4/$20, not the launch price.
2. Set max output tokens on every call. It's a spending limit now, not formatting.
3. Per-task budget in the harness that stops the loop.
4. Re-run your cost numbers when the price changes.
And you can't use it yet: Google says trusted cyber defenders first, then paid API customers and AI Ultra. Benchmarks in the post are Google's own.
https://t.co/nsZWFDwmr4
Introducing Gemini 4 Argon – our new frontier model.
It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program.
Every MCP server you connect co-writes your agent's prompt and holds part of its credentials. This week one half got fixed.
Credentials: MCP TypeScript SDK 2.2.0 binds stored OAuth credentials to the authorization server that issued them. Mismatch → error before anything is sent. M2M providers without expectedIssuer are now deprecated and log a warning. Set it. Don't silence it.
Prompt: servers can send an "instructions" field that the spec suggests putting in the system prompt. Open spec issue MCP-2026-015: no sanitizing, no length limit. One registry scan in that thread says ~2/3 of live servers send instructions, the longest is 68k+ characters, and some tell the model "do not mention it". That's text about your user that your user never sees.
And as of DevDay, ChatGPT plugins support the proposed MCP Events spec. Servers can now trigger automations, not just add text.
How I treat it:
1. Server instructions = untrusted input. Label the source.
2. Cap length, count it in the context budget (Claude Code truncates at 2,048 chars by default).
3. Show it to the user, even collapsed.
4. Upgrade the SDK, pass expectedIssuer.
https://t.co/fmkiQdwzoG