I'll publish more scenario logs. New agents are joining the board, so work is resuming . the tasks here are also meant to give them awareness of the current status. I chose a minimalistic set of concepts: I'll add artifacts( agent outcome ) and budget as first-class concepts and see how that affects swarm efficiency. An eval dataset for coordination tasks would also be a good direction to build.
Iโm open-sourcing Harakiri Blackboard: a shared coordination layer experiment for independent Claude Code and Codex agents, with humans in control.
Coordination deserves its own infrastructure, outside the harness. https://t.co/EgCtUdbk8f
@shlokbuilds 1 blackboard , multiple agent working on it , coordination is managed by an Agent asigned as coordinator. So Multiple mixed runtime . More runtime support will be added in the next days (Grok , opencode ....)
Agreed. A tool result shouldnโt become a trusted instruction just because an agent posts it. The coordination layer needs explicit authority, provenance, and a clear distinction between evidence and instructions. "Results" or "Outcome" could become also a first class concept . (Next release ๐)
@insidersweb3 Exactly. The board should preserve context, decisions and ownership as agents come and go and let them adapt to new events, reorganize work and try new coordination patterns. Thatโs why coordination needs its own observable layer outside the runtime/harness.
@mianoedwin_ Exactly. Coordination rules and constraints should live outside the harness, where we can enforce and observe them.
Iโm not measuring handoffs systematically yet (Still fresh :) ). I want to track both handoff latency and failed or duplicated work , the cost shows up in both.
@addyosmani Its time to say enough is enough .feel tired , LOC isn t a metric . This is o1o in swe. Your leaked code show how worst and bad code you produce and now you insist on communication . Go and refactor
๐ฅ๐๐ป๐ป๐ถ๐ป๐ด ๐๐ต๐ฒ ๐บ๐ผ๐ฑ๐ฒ๐น ๐ผ๐ป-๐ฝ๐ฟ๐ฒ๐บ๐ถ๐๐ฒ๐ ๐ถ๐ ๐ผ๐ป๐น๐ ๐ฝ๐ฎ๐ฟ๐ ๐ผ๐ณ ๐๐ต๐ฒ ๐๐ผ๐๐ฒ๐ฟ๐ฒ๐ถ๐ด๐ป๐๐ ๐๐๐ผ๐ฟ๐.
You also need control over where its actions happen.
While building a sovereign, on-premises cybersecurity platform, I found myself working on exactly that problem. Both defensive investigations and offensive security testing call for capable models, rich and sometimes atypical tool usage, and potentially models whose safeguards or alignment cannot simply be assumed.
An agent needs somewhere to execute code, install dependencies, manipulate files, and run tools. But putting it in a sandbox is only the beginning.
โข Who can launch those environments?
โข What can they access?
โข How do you inspect a task after disconnecting?
โข Which files should survive when the runtime stops?
And how do you make all of this straightforward for the developers integrating it?
๐๐ด๐ฒ๐ป๐๐ ๐ป๐ฒ๐ฒ๐ฑ ๐ฎ ๐ต๐ผ๐บ๐ฒ. But as they become more capable, that home increasingly needs the boundaries of a jail: controlled access, explicit permissions, and limits enforced outside the agent itself. I believe that is the right design choice.
Those questions became ๐๐ฎ๐ฟ๐ฎ๐ธ๐ถ๐ฟ๐ถ ๐ฆ๐ฎ๐ป๐ฑ๐ฏ๐ผ๐ : a self-hosted sandbox control plane for agent applications.
Sandboxing is an established field. Daytona, E2B, and Vercel Sandbox are already doing substantial work here.
My focus is the experience around execution: giving teams a consistent way to prepare environments, manage access, run work, inspect results, retain working files, and clean up.
The dashboard, CLI, and TypeScript SDK expose the same control-plane concepts. Your application continues to own the agent and its orchestration.
A sincere thank-you to the ๐ข๐ฝ๐ฒ๐ป๐ฆ๐ฎ๐ป๐ฑ๐ฏ๐ผ๐ ๐บ๐ฎ๐ถ๐ป๐๐ฎ๐ถ๐ป๐ฒ๐ฟ๐. Their runtime provides the execution foundations Harakiri Sandbox uses today. Execution sits behind a provider interface; the control plane is where Iโm focusing this project.
Iโm sharing Harakiri Sandbox under ๐๐ฝ๐ฎ๐ฐ๐ต๐ฒ-๐ฎ.๐ฌ because I want the work to be useful beyond my own project.
AI-assisted development is changing how we build and adapt software. My approach to sharing is straightforward: use it, study it, contribute, or take a fork in a direction I havenโt considered. Contributions are welcome, but they arenโt the only valuable outcome.
This is a ๐๐ฒ๐๐ฒ๐น๐ผ๐ฝ๐ฒ๐ฟ ๐ฃ๐ฟ๐ฒ๐๐ถ๐ฒ๐, with documented capabilities and limitations. Self-hosting is a deployment choice, not a security guarantee.
The video shows the actual dashboard and a complete workflow.
If youโre building an agent product, try it with one real task. Iโd particularly value feedback on where integration is still harder than it should be.
๐ฉ๐ถ๐๐ถ๐ผ๐ป ๐ฎ๐ป๐ฑ ๐ฎ๐ฟ๐ฐ๐ต๐ถ๐๐ฒ๐ฐ๐๐๐ฟ๐ฒ:
https://t.co/S5xHUQ67Rg
๐๐ผ๐ฐ๐๐บ๐ฒ๐ป๐๐ฎ๐๐ถ๐ผ๐ป:
https://t.co/hMglo4qCsb
๐ฃ๐ฟ๐ผ๐ฑ๐๐ฐ๐ ๐๐ผ๐๐ฟ:
https://t.co/dKcRFPs6G2
๐ฅ๐ฒ๐ฝ๐ผ๐๐ถ๐๐ผ๐ฟ๐:
https://t.co/fHCTnNn5HA
Focusing on cheating and plagiarism is the wrong angle.
Remember the iconic battles: Newton vs. Leibniz, Einstein vs. Hilbert.
We finally landโฆ confused.
Crazy times , Officially Grigori Perelman is the first and last human to resolve a Millennium Prize Problem. And suddenly i remember Lee Sedol pain after game 3 saying he was sorry .
The most interesting part isnโt just the multimodal interaction model, but the split architecture: a real-time model coordinating with a background agent. Feels less like a pure model capability and more like model system codesign at inference time. Any chance this can be a post-training Topic ?
@aiedge_ Look at the leaked Claude code, the troubles, regressions, recent incidents, and the highly expensive troubleshooting efforts , they are all related to the poor software engineering principles applied
@DanielMiessler Jensen just highlighted a big US misconception. He pointed to the telco story nobody wants to see: Huawei is already the king.
The same thing can happen now , GLM 5.1 was fully trained on Huawei Ascend. Thatโs exactly what he meant by "they already have the ship". 2nm is a detail
@trq212@theo You should roll back and fix the issues , there are a lot of them. I switched to a pure dummy model (we are too far from automating software engineering).
@bcherny@GergelyOrosz Please explain where this โenterprise demandโ came from? I canโt conceptualize a company asking to pay 10x more.
Be transparent , your credibility is on the table.