I orchestrate AI agents across tools and projects.
I'm building Session Orchestrator around that work: plan the tasks, run parallel sessions, check the results.
Each session needs a useful handover: what changed, what was checked, what's still open.
@pvncher i’d love a benchmark where the user changes their mind halfway through. “actually, keep the old login flow” at hour two feels like a pretty real test.
@madanmohan16 i'd try a tiny app where they swap jobs halfway through. can the next agent pick up the work from the other two without you explaining it again? that's the bit i'd want to watch.
@vikrantshukla reply triage for me. i've run it alongside existing rules, without letting it change decisions. that gave me actual disagreements to look at before deciding how much to trust it
@0x80Rg i'd happily lose 500 lines of old test output before losing the one sentence where we decided what not to build. the two compaction steps make sense to me