I kept hitting model failures on phones that I couldn't reproduce in the office. LangFuse didn't catch it. LangSmith didn't catch it. So I built the thing I wanted: a flight recorder for AI running on real devices.
https://t.co/JDhV7s4RZK
@irangareddy
Hey — saw your post about the agent wiping the DB / running destructive commands
We’ve built Rykan V specifically for this: policy guardrails that block dangerous shell/file actions before they execute.
You can also run ryk scan to see every dangerous command your agent has run in past sessions.
Open-source here if you want to try or star it: https://t.co/Q1GQ3HqtPZ
@pauliusztin_ next step is being able to control what the agent can do, we're solving that at https://t.co/sucpUzQZrs, free oss. try it out let me know what you think
@arcee_ai Isolated workers help until the episode is a lie. If the worker exits after a half-write, the orchestrator just schedules the next half-write.
@nykdotdev A clean first run is a demo. The graph only counts if a crashed subagent resumes from checkpointed state without rewriting files it already finished.
@tryadaline Named regression buckets are the useful part. Shipping prompt or halt changes against a bug list just churns. Grade the change on the named cluster or you are guessing.