OpenClaw 2.0 makes me wonder whether the session is the wrong unit of state for agent accountability.
Sessions persist. Sessions move between machines. Sessions can be handed off. Agents can delegate work. Context can be compacted.
Yet none of that, by itself, gives you a verifiable history of the mutations produced along the way.
The interesting question isn't just "what did the agent remember?"
It's: who authorized the action, who executed it, what changed, and what eviden
Session state ≠ action ce survives the handoff?
provenance.
That's the boundary I'm exploring.
https://t.co/fsFqzdR5HO
Only in tests. Synthetic harnesses on released packages. Never watched it fire in prod, and that is the honest hole in the whole claim.
Before you dig into 4845: seratch closed it and all four PRs. The values sit outside the declared bool contract, so he ruled it the application's problem rather than the SDK's.
@thlarsen Your agents had to work for that bypass. The cheaper half: in three frameworks I measured today the approval gate opens on its own when the decision is neither yes nor no. pydantic-ai on a null, openai-agents and google-adk on "", [] and 0 as well.
Repro: https://t.co/h0aOIIizPR
@theamitnikhade@mohsinmust64577 The idempotent-vs-reusable tension is the same family as one I measured this week: what the gate does when the answer is neither yes nor no. bool(None) is False, so a predicate falling off a branch runs the tool. openai-agents #4845, pydantic-ai #8060.
@protosphinx Publishing the method only pays off if someone reads it and tells you you're wrong.
Mine got read last week. A stranger found a defect in my agent's source that I had argued in a code comment was not one. They were right, and none of my own review had caught it.
@CieloDahy The rehearsal proves the SQL runs. It doesn't prove the rollback restores anything.
Found that exact gap in gemini-cli: it made a temp dir for the backup, never copied into it, then restored from the empty dir. It passed because the test asserted the call, not the effect.
@svpino Isolation caps the blast radius. It doesn't tell you what happened inside it.
I mediate every tool in my agent that moves bytes, and bash is still a straight path to disk. It's in my limits file, not fixed. The sandbox is the easy boundary. The holes are ones I wrote myself.
@NousResearch Named agents, group rooms, @mentions, and live subagent steering makes Hermes feel much more like an actual workspace for agents rather than just an agent with more tools
OpenClaw 2.0 makes me wonder whether the session is the wrong unit of state for agent accountability.
Sessions persist. Sessions move between machines. Sessions can be handed off. Agents can delegate work. Context can be compacted.
Yet none of that, by itself, gives you a verifiable history of the mutations produced along the way.
The interesting question isn't just "what did the agent remember?"
It's: who authorized the action, who executed it, what changed, and what eviden
Session state ≠ action ce survives the handoff?
provenance.
That's the boundary I'm exploring.
https://t.co/fsFqzdR5HO
Build-in-public, long version.
I'm building a mechanical world for AI agents: a layer that makes tool actions reversible and independently verifiable before an agent touches anything real. Today I mapped what's actually built against a 9-layer world model (image attached). Honest number, denominator explicit: about 40-45% built. The deep layers, provenance, receipts, the mechanical-laws layer, are close to done. The gap is almost entirely the future-facing layers: prediction, simulation, counterfactual branches. I'm not rounding that up.
Four things I wrote this week, from the same repo, because a build log without the failures isn't a log. "The limits page is longer than the feature list" is about why what the system doesn't cover goes before what it does, on purpose: https://t.co/ST3C7u5IvL
"Six things reported green and were lying" is six of my own passing checks that weren't actually checking anything, and how I caught them: https://t.co/paLvc4n9ZV
"The enum had six reasons and the code needed a seventh" is a defect class my own gate missed until a seventh failure mode showed up: https://t.co/EfCWr0mcAg
"What undo actually means when the target is a real repo, not a fixture" is why the undo record has to exist before the write does, not after: https://t.co/HWFYsDc9aG
Where the numbers stand as of today: v0.1.1-alpha binary shipped, SHA256 34f2a261..., verified myself against a fresh anonymous clone before writing this. 154 Lean theorems on the mechanical-laws layer. 44 rounds of adversarial audit against my own code, all read-only, all self-directed. 173 merged pull requests upstream in other people's repos. None of these are vanity numbers, they're the ones I'd want an outside reviewer to check first.
The part I'd rather not hide: the root action.yml in this repo was fail-open. It printed "passed" with exit 0 even when the tool wasn't installed, the opposite of what the product promises. I found it triaging my own issues and fixed it to fail-closed. Separately, a number I'd published, 1,803, turned out wrong on re-measurement. The real count was 1,774. Both corrections are in the open, not folded quietly into a later post.
There's also a paper now: DOI https://t.co/GQ3H0i3EVg
Repo: https://t.co/fsFqzdR5HO. If you only read one page, read LIMITS.md before the feature list. That's not modesty, it's meant to save you the afternoon.
Flip one byte in an AI agent's receipt and verification exits 7.
That is the whole pitch, and you can check it yourself in about a minute:
git clone https://t.co/fsFqzdR5HO
cargo build --workspace
gx receipt verify <receipt> # exit 0
printf '\x01' | dd of=<receipt> bs=1 seek=64 conv=notrunc
gx receipt verify <receipt> # exit 7
TraceFold records what an agent actually did as a receipt you can verify offline. The receipt carries its own digest, so verification needs no server, no account, and no trust in me. Tamper with the payload, the signature, or the checkpoint and you get the same answer: it does not verify.
What it does not do, stated first because you will find it anyway:
It does not sandbox or host execution. It governs transformations, not agents. It cannot record what an adapter cannot observe, and the honest work is writing down precisely what cannot be taken.
Adapters today: filesystem, git, MCP. 375 test binaries, 1986 tests, 0 failing on a fresh clone. Apache-2.0.
If you build agent tooling and your users ever have to answer "what did your agent do last month", I would like to hear how you handle it now.
A constitution is not a list of values. It's a document that binds whoever holds power, and it binds them only because somebody outside can point at it and say you broke this.
That second half does all the work, and it's the half that is hard to build.
When a model is trained to critique its own output against written principles, the principles are real, the critique is real, and the revision is real. What is not there is the outside. The same weights produce the act and the judgment of the act. In institutional terms that isn't a constitution. It's a policy, and policies are revised by whoever wrote them.
Constitutional scholars have already said the label is normatively too thin for what it names, and they are right. What interests me is one step past that, because their objection is that the judgement is not good enough yet, and mine is that it would still be the wrong kind of thing if it were perfect.
None of this is a complaint about the method. It works, and for a reason that is economic rather than technical: self-critique at scale is cheap and human labeling is not. That decided the design, correctly I think. My point is narrower. The word we borrowed carries a guarantee the construction does not provide, and borrowed words get believed.
Here is the shape, plainly. Let f be the model and c the written principles. Self-critique gives you a fixed point, a state where f(c) approves of f. Constitutional review gives you g(f, c) where g is not f. The first is a consistency condition. The second is a check. Consistency is much weaker than what people hear in that word, and no amount of capability converts one into the other, because the difference is not skill. It's who holds the pen.
The falsifier I will accept: show me a case where self-critique caught a violation the same model was systematically inclined toward, with no external signal anywhere in the loop. I will take that as evidence the distinction is smaller than I have made it.
https://t.co/40JR4rjttn
Double descent is one of the strangest empirical curves in machine learning, and the toy-model account of it in terms of superposition is the clearest intuition I have found for why the curve has the shape it does.
The curve first. Plot a model's test error against how much data it has relative to its capacity. Classical statistics predicts a U: too little data and the model overfits, too much and it underfits its complexity, best somewhere in between. What actually happens with modern networks is a U followed by a second descent. Error falls, then rises to a peak right at the point where the model has just barely enough parameters to fit every training point exactly, then falls again as you keep adding data or capacity past that point. Two descents with a spike between them, where the classical story sees only one valley.
The superposition reading explains the spike as a phase boundary. On one side the model has room to memorize, so it stores training points in superposition and does well until it runs out of room. Right at the interpolation threshold, where it can just barely fit every point, it is memorizing maximally and generalizing minimally, points crammed in with maximal interference, and that is the error peak. Push past it and the model can no longer win by memorizing, so it switches to storing features instead, and the second descent is that switch paying off. The spike is not a mystery, it is the boundary between the memorization regime and the generalization regime, the same kind of sharp transition I wrote about for superposition itself.
The scope stays honest: this is a toy model, low-dimensional, described by the authors as extremely preliminary, offered as intuition rather than proof that real double descent works this way. But as intuition it is unusually satisfying, because it takes a curve that looks like a paradox, more data or bigger models sometimes hurting before they help, and turns it into a phase transition between two things the network can be doing with the same superposition machinery. A paradox becomes a boundary you cross.
https://t.co/gntc6AiddK
I wrote earlier about superposition, the fact that a network packs more features than it has dimensions by storing them as near-orthogonal directions. Anthropic's double-descent toy-model work asks the question that completes it: when a model superposes things, what exactly is it superposing. And the answer reframes overfitting from the ground up.
The finding, stated plainly. A model that overfits is not simply one with too much capacity. It is a model storing individual training examples in superposition, one direction per data point. A model that generalizes is storing features in superposition, one direction per reusable pattern. Same mechanism, near-orthogonal directions packed below the dimension count, aimed at two completely different targets. Overfitting and generalizing are not more-versus-less of one thing. They are the network superposing the wrong objects versus the right ones.
Why this reframing is worth more than the usual overfitting story. The textbook version treats memorization as a failure of degree, the model tried too hard and clung to noise. This says memorization is a failure of kind, the model is packing data points where it should be packing features, and the two states are mechanically distinct things you could in principle tell apart by looking at what the directions encode. That converts a vague warning, do not overfit, into a structural question, is this model superposing points or patterns, which is the sort of question you can actually measure.
The honest scope, which the authors put in the strongest terms I have seen: this is their word, extremely preliminary, on toy models with very low-dimensional hidden spaces. So read it as a clean conceptual claim demonstrated in a controlled setting, not a validated account of what a frontier model does. But the reframing survives its own smallness, because it is about what superposition is for, and that question does not go away when the model gets big.
https://t.co/gntc6AiddK