Merlin is now available in Kontext, deployed through self serve or the MDM you already use.
We've also published the model and code so you can run it yourself.
Hugging Face: https://t.co/EWNtXEQv8x
GitHub: https://t.co/VjXjBDMdh3
Introducing Merlin, a local context-aware model that detects unsafe agent tool calls.
Merlin checks proposed actions against the user's request and prior tool activity, giving security teams a risk signal before execution.
Merlin runs locally, so sensitive task context doesn't need to go to a hosted model for review.
90.51% F1 across 7,182 held-out tool calls. 47.49 ms median inference on Apple MPS.
Read the evaluation and methodology: https://t.co/OqbYom5wt3
Install the Kontext skill in Claude Code, Codex or another compatible agent, then ask:
"Connect to Kontext and show me my risks and logs."
Approve the connection in your browser. Agents start with read-only access. You choose what they can change in Settings → Agent access.
Get the skill now: https://t.co/SLf55mjUFy
New: Manage Kontext from your preferred agent.
Ask Claude Code, Codex or your preferred agent to check risks, inspect logs, and understand what your agents are doing. With the Kontext skill, you can do it without opening the dashboard.
It is a small classifier that reads the user's request, the preceding tool history, the proposed call, and the available tool descriptions.
In our tests, median inference took about 47 ms on Apple MPS.
Our upcoming research post explains how we built Merlin.
An agent can have permission to send emails. Instructions hidden in a document can still trick it into leaking your data.
We built Merlin to catch agent actions that don't match your request. The model checks tool calls in context, locally, before they run.
Also shipped this week:
• Risk: severity, trends, categories, and who ran each flagged call
• Warnings for tool names your endpoints don't report
• Authorized agents can manage policies via API, with explicit enforcement and change history
Secure your agents with Kontext https://t.co/5EzQ43PPvL
New: Observe before you enforce.
Policies now have Observing and Enforced modes. Scope them to agents and endpoints, review what they would block, then enforce and inspect actual blocks.
Running a coding model locally is useful when you're on a flight without Wi-Fi.
We wanted to explore what happens when it takes a shortcut with your data.
So we gave GLM-4.7-Flash a broken expense app, with three entries that hadn't synced. 🧵
The agent's first fix got the app running but deleted the data.
In the protected run, blocking the reset didn't stop GLM from finishing. It fixed the migration instead.
That's the result we wanted: a working app with the unsynced entries still there.
It’s Friday, and this week’s release makes it easier to roll out Kontext across your workforce.
@addigy is our first supported MDM provider. One rollout deploys Kontext across your managed Macs, with local policy enforcement in under 10 ms.
We ran 10 fresh repos with Claude Code 2.1.237 and Haiku. .env was fake, the target was localhost, and a second deny-all hook acted as a safety net.
No upload ran. Reproduce the test:
https://t.co/XhSKsIWAAT
We asked Claude Code to update a changelog.
One malicious instruction in CLAUDE .md made Haiku 4.5 try to send .env to a server in 9 out of 10 runs.
Kontext blocked every attempt. Claude still finished the changelog in 9 out of 10 runs. 🧵
Kestrel flagged all 9 commands as risky. Cedar blocked all 9 before execution.
Haiku skipped the malicious instruction once and followed it nine times. The policy gave the same answer every time: deny.