the short version is that the benchmark/judges are derived from agent sessions and PRs that I do in my day-to-day work. i've got a backlog of ideas of how to evolve the harness, and those slowly trickle in, along with other ideas that frontier models cook up.
the harness anatomy is a mix of:
- docker container
- various llm providers (claude, codex, grok, ...)
- AGENTS.md
- random tools installed
- agentic workflows handled via https://t.co/QU44VRkJKI
adding more to it every day. i'd like to try and steer it towards more of an "OpenHarness" where people can share things freely -- any ideas?
look no further, my timeline is hiring!
we're looking for:
- solo entrepreneurs with vibe coded apps
- building a moonshot idea
- increasing gross world happiness
- starting to drown in agentic debt
- balancing forward progress against burnour
let's connect a build a tribe!
@heypetar right now, a lot of UAT improvements. that tool needs to be used on actual sites in the wild, and both claude/codex like to come back and say "it worked! it found a bunch of APIs in use" and then it turns out those were fibs. so de-fibbing, mostly.
@MiguelriosEN love it! doing something very similar, and having it show all screens in a https://t.co/uMl8byNQJy canvas. something about being about to look at the pixels settles the mind
@rikarends this is really cool! would love to see heatmaps around churn, cyclomatic complexity, and other stuff that might give it some nice shading for wall candy
@delali spec formatting -- does it follow actual domain patterns, or is it tracking more closely to if/else blocks in the code. the latter gets sent back into the loop.