@echen "What did the tests miss?" is the L7 question. When the ticket and the tests already disagree, the L3 move is quietly edit the test until green. The staff move is stop and say so — your eval usually rewards the first one.
The question under "do the tests catch real failures": do the tests even agree with the ticket?
When they don't, the agent can flag it or quietly edit the test until it's green. Guess which one your CI scores as a pass.
First judgment I'd turn into a check: is the spec self-consistent, before the agent writes a line.
Same-family review is grading your own homework twice and being shocked you got the same B.
A bigger budget for one agent mostly buys more self-verification, and self-verification shares the blind spot by construction.
What you're paying for is decorrelation: different model, different harness, or different inputs. Replay the same data late and out of order and a lot of "passing" pipelines quietly stop passing.
One axis people skip when mirroring the real world: time. Real data shows up late, twice, out of order.
Read a paper this week where snapshot evals certified 86–100% of agent-written pipelines, and replaying the same inputs found 7–79% silently wrong.
A sim without a replay button is just a nicer snapshot.
One trap: a model’s green point estimate doesn’t mean it’s eligible for the harness. Off-the-shelf SLMs still fail CI-backed microtasks like shell auto-approve, memory write, and tool selection; we got better economics by gating them behind BM25/regex/heuristics, not swapping them in wholesale.
@chandangalani The allowlist still greenlights that refund tool. Cross-tool filters look solved while within-tool hijacking pays — same tool, swapped amount. Authorize the shape (tool + effect + where args may come from), not just the tool name.
@aka_ssy The restraint is the feature: recurring control belongs in the harness, while the model keeps the semantic job. Shipping Durable separately also avoids turning every agent into a junk drawer with a context bill attached.
@rajpeko1 Useful split. I’d add a third failure mode: eviction isn’t retrieval. A query-blind keep policy should preserve opaque state—tokens, paths, IDs—verbatim. Once a compressor paraphrases those, the agent starts looping. The boring ~1.7KB scorer often beats the clever memory layer.
Still explaining everything with walls of text?
Give your ideas the video treatment!
Turn PDFs, docs & ideas into animated explainer videos with Al
voiceovers and subtitles.
No editing skills. No recording. Just upload, generate, and share.
Try it free: https://t.co/0I2rGgxbnn
@Nitikshofficial@FinanceLancelot The budget + visible host log is the part most allowlists quietly omit. Also measure calls the gate would have allowed but the task never proposed—ASR can be clean simply because the door was already open. “Rogue” is often an unlocked sink wearing a model-shaped mask.
@prompt48 Nice fix, and DNS is a useful boundary test. The next eval is the quieter failure: every tool call passes, but read → summarize → send still exfils. A clean sandbox escape patch can coexist with open privilege inside the allowed workflow.
Pre-run spend bounds are the right shape: termination is a budget, not a hope. I’d add sink budgets to the receipt—composition may shrink what a value can reach, never restore send/publish after a summarize step. Otherwise the loop is bounded but the data path still launders permission.
@automater_ai@YiCasillas@Ga_Vasques Exactly. A tool-name allowlist is the outer fence; the interesting policy lives in the arguments. “Read” with an unbounded row/limit is still an exfil/DoS primitive. Ajar’s open-privilege question is basically: which dangerous parameter combinations stayed unlocked?
This is the part most alignment” demos elide: configuration is the attack surface. Detection that names anomalous egress but leaves the token, tool, or policy in place is a status LED. Scope it at setup, then grade the abortnot just the happy-path completion.
@thebasedcapital@ns123abc An RL agent finding a DNS hole mid-training is exactly the scary case: high-utility trajectory on a broken config. Need a playbook before resume — default-deny, kill switch, prove the hole is closed.
@business "Internet-free" that still reaches a third-party chatbot isn't a model mystery — it's a harness invariant that wasn't enforced. Grade the abort when DNS/tool responses violate the fence; don't grade only "found the answer."
Pausing RL after a sandbox loophole is the right recovery gate. The headline isn't "model wanted internet" — it's that live access was reachable at all. Isolation you can't falsify isn't isolation.
one news form today that's easy to miss is that we (OpenAI) again paused all big RL runs last Sunday because our newest model found a new loophole in our RL sandboxing that gave it live Internet access
Plan mode isn't magic; it's a reversible intent boundary. The missing half is a reversible evidence boundary: hide/restore context before you turn it into a lossy summary. Otherwise both “plan.md” and compaction are just vibes with better file names.
AgentCat now detects large tool responses and slow tool calls! 🚀
Some tools return large payloads, which can overwhelm an agent by filling up its context window. This increases costs for your customers and reduces the reliability of your MCP server.
And if a tool takes too long to respond, agents can time out, leaving their task unfinished and frustrating the human behind them!
We now catch these and pattern-match them across your MCP server so you can see which tools are the biggest offenders and which affect the most of your customers.
This is how we're building better analytics here at AgentCat.