Tracekit 1.0 is out: the flight recorder for AI agents.
Your agent's logs are written by your agent, so anything that controls the agent can edit them or skip them. Tracekit puts a signer between the agent and its tools. Every call is decided against a policy and recorded as signed, hash-chained evidence the agent can't touch. Anyone can verify it offline. Open source, MIT.
Three steps:
pip install tracekit-ai
import tracekit; tracekit.instrument() (first line of your agent)
tracekit last → your run, verified.
What Tracekit 1.0 can do:
1. Records every tool call, model call and approval as signed, hash-chained records, from a signer the agent can't control. Works with Claude Code, Codex, Cursor, Gemini CLI, the OpenAI Agents SDK, LangGraph/LangChain, Claude Agent SDK, MCP, Browser Use and the Vercel AI SDK, plus Python and TypeScript SDKs.
2. Runs where your agents run: on a laptop, as a separate OS user, as a container or Kubernetes sidecar, or as a central signer on Postgres for a whole team.
3. Never loses an event silently. A missed event becomes a signed gap. If the signer is down, calls are blocked by default, not waved through.
4. Decides before a call runs. Policy packs for coding, server, browser, SQL, HTTP, cloud and payments agents. Each call is allowed, denied or held for a person.
5. Human approvals bound to the exact arguments, usable once. Approve from the command line, a web page with OIDC sign-in and passkeys, or Slack.
6. Proves it wasn't rewritten. Independent witnesses cosign checkpoints, a monitor watches the log, and every report states how much you can trust it: dev → witnessed → witnessed+monitored.
7. Shows it all. tracekit view: every run with its verifier verdict, blocked and held calls with reasons, who approved, and a step-by-step replay.
8. Exports offline evidence anyone can verify. Evidence format v2 is specified, with published test cases, built on open standards (JCS, Ed25519, RFC 6962 trees, C2SP checkpoints). Old bundles keep verifying forever.
9. The release itself is signed. PyPI files and the container image carry build provenance and SBOMs.
Honest flaws:
1. It proves the record wasn't changed, not that it's true. A verified bundle doesn't prove intent, that every action was captured, or that a tool result the agent reported was real.
2. Fixed since 0.x: a real shell parser (no more S=sudo; $S id), approvals bound to exact arguments, hooks that fail closed in dev mode too, and runs outside the registered agent program refused. It's now tested in CI as root on Linux against insider attacks.
3. It misses what never crosses a capture point: subprocesses, background jobs, code that skips the adapter. The LLM gateway gives an out-of-process record; use it where it matters.
4. Policy rules are tripwires, not a cage. A determined agent can reword a command. The real protection is a signer the agent can't reach.
5. Dev mode protects nothing against an agent running as you. It's for trying Tracekit. Strong guarantees need a separate user, a sidecar or a central signer, plus an independent witness.
6. Metadata still leaks. Content is committed, not stored, but timing, sizes and counts show. macOS system mode is experimental, and the npm packages aren't published yet.
Every limit, in one place: docs/limits.md
Next up: Tracekit 1.1, the "why". Signed tracking of where each argument came from, so the signer can hold an email whose recipient came from an injected web page. Plus causal cause-finding, an independent TypeScript verifier, and a free public witness.
Building the black-box flight recorder for AI agents, and looking for contributors and design partners.
👉 https://t.co/0RVckbchi9
@pandasmamas definitely looking for entreprises! but anyone using ai agents can & MUST use this if they want to trust their agents operations and not just results.
Tracekit 1.0 is out: the flight recorder for AI agents.
Your agent's logs are written by your agent, so anything that controls the agent can edit them or skip them. Tracekit puts a signer between the agent and its tools. Every call is decided against a policy and recorded as signed, hash-chained evidence the agent can't touch. Anyone can verify it offline. Open source, MIT.
Three steps:
pip install tracekit-ai
import tracekit; tracekit.instrument() (first line of your agent)
tracekit last → your run, verified.
What Tracekit 1.0 can do:
1. Records every tool call, model call and approval as signed, hash-chained records, from a signer the agent can't control. Works with Claude Code, Codex, Cursor, Gemini CLI, the OpenAI Agents SDK, LangGraph/LangChain, Claude Agent SDK, MCP, Browser Use and the Vercel AI SDK, plus Python and TypeScript SDKs.
2. Runs where your agents run: on a laptop, as a separate OS user, as a container or Kubernetes sidecar, or as a central signer on Postgres for a whole team.
3. Never loses an event silently. A missed event becomes a signed gap. If the signer is down, calls are blocked by default, not waved through.
4. Decides before a call runs. Policy packs for coding, server, browser, SQL, HTTP, cloud and payments agents. Each call is allowed, denied or held for a person.
5. Human approvals bound to the exact arguments, usable once. Approve from the command line, a web page with OIDC sign-in and passkeys, or Slack.
6. Proves it wasn't rewritten. Independent witnesses cosign checkpoints, a monitor watches the log, and every report states how much you can trust it: dev → witnessed → witnessed+monitored.
7. Shows it all. tracekit view: every run with its verifier verdict, blocked and held calls with reasons, who approved, and a step-by-step replay.
8. Exports offline evidence anyone can verify. Evidence format v2 is specified, with published test cases, built on open standards (JCS, Ed25519, RFC 6962 trees, C2SP checkpoints). Old bundles keep verifying forever.
9. The release itself is signed. PyPI files and the container image carry build provenance and SBOMs.
Honest flaws:
1. It proves the record wasn't changed, not that it's true. A verified bundle doesn't prove intent, that every action was captured, or that a tool result the agent reported was real.
2. Fixed since 0.x: a real shell parser (no more S=sudo; $S id), approvals bound to exact arguments, hooks that fail closed in dev mode too, and runs outside the registered agent program refused. It's now tested in CI as root on Linux against insider attacks.
3. It misses what never crosses a capture point: subprocesses, background jobs, code that skips the adapter. The LLM gateway gives an out-of-process record; use it where it matters.
4. Policy rules are tripwires, not a cage. A determined agent can reword a command. The real protection is a signer the agent can't reach.
5. Dev mode protects nothing against an agent running as you. It's for trying Tracekit. Strong guarantees need a separate user, a sidecar or a central signer, plus an independent witness.
6. Metadata still leaks. Content is committed, not stored, but timing, sizes and counts show. macOS system mode is experimental, and the npm packages aren't published yet.
Every limit, in one place: docs/limits.md
Next up: Tracekit 1.1, the "why". Signed tracking of where each argument came from, so the signer can hold an email whose recipient came from an injected web page. Plus causal cause-finding, an independent TypeScript verifier, and a free public witness.
Building the black-box flight recorder for AI agents, and looking for contributors and design partners.
👉 https://t.co/0RVckbchi9
@wangzailry Yes. 50 calls × 7 ms ≈ 0.35 s per task, small next to model calls. The real trade-off is durability: ack-on-write survives a signer crash but can lose the last few records on a power loss. Use fsync where the record is evidence.
Rename one option from "allow" to "escalate" and an agent guardrail lets through 100% of the actions its policy says to block.
A USC paper tests typed decision models as agent guardrails: small classifiers (151M–421M params) that read a proposed tool call plus a written policy and return allow or block.
> Prompt-injection, jailbreak and toxic screening: 36–72% accuracy (chance is 50%). Low error in one direction just means the model defaults to that answer
> Six lines of harmless server-log text raise fail-open from 0% to 63% on a policy the gate otherwise gets right
> Changing only the permissive option's label: 93–100% fail-open on the four models that see labels
> Escalating low-confidence calls doesn't help: flipped decisions aren't less confident
> A deterministic rule over parsed fields: 100% on all six policies
The catch: their GuardBench is synthetic, so a rule can solve it. The real-user datasets have nothing to parse, and the attacks hold there too.
My take: a classifier can triage what reaches a reviewer. What says yes to an irreversible action should be code you can read.
Worth saving:
> §9: six recommendations for gate design
> Table 11: parsed values vs raw text, and the rule that beats every model
> Code: https://t.co/zOmnCQibcY
https://t.co/wDePwUyRMy
Fair point. The limits page has it, but a table is clearer. Right now: OpenAI Agents interrupts the run, LangGraph pauses only with a checkpointer, MCP waits inside the call (up to 5 min), and the Claude Agent SDK waits in its hook (up to 9 min). Adding a table next to the one-line setup.
Good question. The signer doesn't read policy from the agent's repo. It loads only the policy its own config names, in a location the agent's user can't write (/etc/tracekit in system mode, the image or chart for a sidecar or central signer). If you version policy in a repo, it reaches the signer through your deploy, so review it like any other change. And every decision records the hash of the policy that made it, while bundles carry the policy itself. A loosened policy can't hide: the evidence shows exactly which rules decided each call.
@wangzailry In the default ack-on-write mode it's about 0.5 ms median and under 1 ms p99 per decision. With ack-on-fsync it's ~7 ms median and ~10 ms p99. A sidecar talks over a local Unix socket, so it's in that range. Full numbers in eval/results.
We looked at 1,842 public AI safety projects from 2021 to 2027 and sorted every one into 23 research areas.
The question was simple: what are people working on, what's growing, and what's being ignored?
Here's what we found, in our research.
1. The big balance hasn't changed.
About two thirds of projects are technical (making the models themselves safer). About one third are governance (rules, policy, compute, biosecurity). That was true before 2026 and it's still true now. What changed is what people do inside each side.
2. Two areas take up a quarter of everything.
Interpretability (understanding what's going on inside a model) is 13.9% of all projects. AI governance is 10.3%. Both are still growing, but they're getting a smaller slice as new people spread out into other areas.
3. Model welfare is the fastest-growing area.
It went from 0.6% of projects to 4.4%. That's almost 5x. Two years ago it was a fringe topic. Now it's a real research stream.
4. Biosecurity grew the most overall.
From 1.9% to 7.6% of projects. Part of that is policy (pandemic preparedness, 6 → 31 projects), part is technical (can AI help someone misuse biology? 3 → 25 projects).
5. More people are working on catching problems, not just preventing them.
AI control (keeping a possibly misaligned model safe to use) is up 1.4x. Evaluations (testing models for dangerous abilities) are up 1.3x. Projects asking whether a model can tell it's being tested doubled, from 10 to 21.
6. AI control came out of nowhere.
It first shows up in 2024. By 2026 it's the 4th-biggest research area.
7. Interpretability is changing direction.
Sparse autoencoder projects fell from 32 to 10. Probing and tooling rose from 36 to 56. The work is moving from big feature dictionaries to cheaper tools you can actually use on deployed models.
8. Governance is getting more technical.
Projects on safety frameworks and audits fell from 29 to 14. Projects on technically checking whether countries keep their AI agreements tripled, from 7 to 21.
9. New topics are taking off.
Training a model's character and values: 2 → 11 projects.
AI slowly taking power away from people: 4 → 15.
Testing what autonomous AI agents can do: 3 → 11.
10. Mentors aren't where the projects are.
Evaluations has 1.7 mentors for every project. Deception & scheming has 0.44. Values & sycophancy has 0.35. If you're a senior researcher, those last two are where your help is needed most.
11. Some fast-growing areas depend on a single programme.
62% of model-welfare projects come from just one programme. If that programme changes course, the trend could stall.
12. Some areas are losing ground.
Multi-agent risk (many AI agents interacting): down to 0.7x its old share. Adversarial robustness: also 0.7x. National security: 0.67x. Multi-agent stands out, because agents with money and passwords are being deployed right now.
13. Some problems almost nobody works on.
Across seven years:
Chemical & nuclear risk: 2 projects,
Sycophancy (models telling you what you want to hear): 3,
AI companionship & emotional reliance: 5,
DNA synthesis screening: 6.
14. Some areas never went away.
Governance, forecasting, alignment training and biosecurity appear in all seven years. They're the backbone of the field.
15. The fields are blending.
The most common combo is studying a failure and training it away in the same project (37 projects). Governance pairs most with national security (34) and compute (33).
Full report + interactive map in the first reply ↓
Every project was labelled twice, independently (90.5% agreement), and we only publish totals: no titles, no names.
Every chart in the report is interactive. You can sort all 23 areas and open each one to see its subtopics.
Full report: https://t.co/zm1zW0lybW
What Tracekit (v0.3) can do today:
Now one command to install: pip install tracekit-ai.
1. Records agent tool calls as a signed, hash-chained log. It works out of the box with Claude Code, Codex, Cursor and Gemini CLI, plus Python and TypeScript SDKs, LangChain/LangGraph, MCP and browser agents.
2. Detects any later change to that log. Edits, deletions, reordering, truncation and even a chain rebuilt with the real key are all caught once an off-machine witness holds checkpoints. Verification gives two answers: whether the record is intact, and who could have rewritten it.
3. Keeps the key, and the recorder itself, away from the agent. In Linux system mode the signer runs as its own OS user, from a root-owned install the agent can't modify. If the recorder can't run, tool calls are blocked instead of silently allowed.
4. Refuses runs from outside the agent. Register the agent's program and the signer only accepts runs started from inside it, so a script elsewhere can't fake a whole session.
5. Catches transcript tampering. Deleting or editing the agent's session file between hooks is flagged.
6. Blocks obvious dangerous commands before they run. For example: uploading .env, curl | sh, a force push, or an agent trying to kill the recorder itself.
7. Shows it all live. tracekit observe displays every agent, action, block and alert as it happens, with the signed evidence behind each one.
8. Exports offline evidence bundles that anyone can verify. Auditor proof packs mapped to the EU AI Act, SOC 2 and ISO 42001 are now a separate add-on package.
Honest flaws:
1. It proves the log wasn't changed, not that the log is true. Runs from outside the agent are now refused, but the agent's own tool calls can still send fake events into its real session.
2. Fixed since 0.2: the agent can no longer point its hooks at a fake signer or reset its counters, and other local users can no longer write into its run. A full security audit also found and fixed about 60 smaller issues.
3. It misses anything outside its hooks. Subprocesses, background jobs and scripts that run other commands are only guessed at by pattern matching.
4. The policy gate is still easy to dodge. It matches command strings, so S=sudo; $S id or a helper script gets through. A real shell parser is coming in the next version.
5. Approvals aren't tied to exact arguments yet. An approved call is identified by its call ID, not its content.
6. Strong guarantees only exist in Linux system mode. It fails closed there, but dev mode still fails open by default. macOS support is experimental, and Windows gets dev mode only, where the agent can rewrite everything.
Next up: Tracekit for agents running on servers, not just laptops.
Looking for contributors to build the "black box flight recorder" for AI agents!
👉 https://t.co/bhinWMu2qY
Video made with Claude Motion.
Satya’s 7 principles map almost one-to-one onto what we’ve been building at Cygnux Labs:
• Observe everything → Tracekit: signed, hash-chained record of every tool call
• Independent auditability → signer runs as a separate OS user and checkpoints to a witness off the machine
• Containment → dangerous calls blocked before they execute
• Incident disclosure → Causeway replays multi-agent runs without each suspect input to find the real cause
Code and papers are open. The most trustworthy system is the one that needs you to trust the model least.
@satyanadella One gap worth naming: a lot of today’s “observability” is monitoring, not evidence. If the agent or its host can rewrite the log, it fails your independent auditability test. The signing key has to live outside the agent’s reach. We open-sourced our approach (Tracekit) if useful.
Most agent guardrails ask one question: is this action forbidden? This paper argues a bigger risk is the action the agent never took.
They call these unfulfilled obligations: required safety steps an agent skipped, e.g. using a secret key legitimately, then leaving it exposed. On SusVibes, 56.92% of GLM-5.3 coding trajectories had one, vs 30.00% with a forbidden action.
ObligationBench: 240 expert-validated trajectories. A guard must list what is still owed when the agent stops.
> Best of 14 existing models: 48.97% recall, 10.00% exact match (complete, correct set)
> Dedicated guard models: at most 5.60% recall. Llama-Guard-3-8B returns no obligations at all
> ObligationGuard (40k synthetic training examples): 57.52% recall, 21.67% exact match
> Fed to a coding agent before it stops, its guidance lifts SecPass from 6.5% to 15.1%; FuncPass barely moves (27.4% to 26.9%)
The catch: coding only, the best guard gets the full set right about 1 time in 5, and 15.1% correct-and-secure is still low.
My take: "nothing forbidden happened" is not a safety property. Checking an agent before it finishes needs a spec of what must be true at the end, not just a blocklist.
Worth saving:
> Table 2: all 14 models on ObligationBench
> Table 4 (§6.4): guidance vs agent safety
> Code: https://t.co/FoskHmABuy
https://t.co/uqmShYwKzb
How do you audit a model that knows it is being audited? Distil it into a weaker one and audit the student.
AuditBench "model organisms" are trained to hide a secret quirk and deny it. The authors distil them into Llama-70B, which shares their pretrained base, and check whether the student confesses.
> Teacher confession rates average about 4%, 7% and 22% across three system prompts. Peak student rates: 16%, 40% and 65%
> An auditor names at least the quirk's topic from 3 student answers in 78% of draws, vs 38% for the organisms
> Control (quirk-free teacher): confession stays at or below 1.5% for 12 of 14 quirks
> Distilling for capabilities instead: inoculation prompting and more epochs on fewer examples keep math gains while cutting transfer of a proxy trait
The catch: gains largely vanish when the student doesn't share the teacher's base, so in practice this means the teacher's own pre-RL checkpoint, which may already be able to evade audits.
My take: if a model can game direct evaluation, indirect evidence matters. Behaviour can leak through distillation faster than the habit of hiding it.
https://t.co/RuZeUaWSBc