A completely buried aspect of this paper.
It's not about building a single harness.
They build K harnesses
Each to handle different clusters of tasks, and only evaluate edits on harness_i within cluster_i
The harness variant ensembling is where the gains are
First refusal I've ever gotten from Anthropic (claude code).
Apparently generating descriptions of waste streams that involve anthrax is a no-no
I guess my eval set just won't contain BSL-1 agents...
@plugberryox@amytam01 I am on the exact opposite side of this.
It's never been easier to try things, fail, and use the learnings from that failure to refine your taste.
The people who leverage AI properly will become super-tasters.
We just open-sourced Paperclip: the orchestration layer for zero-human companies
It's everything you need to run an autonomous business: org charts, goal alignment, task ownership, budgets, agent templates
Just run `npx paperclipai onboard`
https://t.co/wuDdEmrSMx
More 👇
If anyone's still looking for this, here's a plugin:
1. Auto updates docs based on code diffs
2. Enforces progressive disclosure of repo context
3. Plans as first class citizens
4. Migrates existing docs on init
PRs welcome
https://t.co/aPRJcu7LC0
It feels like someone should make a post-git-hook where it asks the AI model to look at the diff of what you changed for a merged PR and update the repo’s various readmes and other documentation to make it easier for an LLM to be able to write code and reference things faster rather than reading every single line of source code that might be relevant constantly. The agents need their own docs.
@Suhail If anyone's still looking for this, here's a plugin:
1. Auto updates docs based on code diffs
2. Enforces progressive disclosure of repo context
3. Plans as first class citizens
4. Migrates existing docs to recommended architecture on init
PRs welcome
https://t.co/aPRJcu7LC0
The insight is deceptively simple:
The bottleneck in AI development isn't AI capability.
It's human clarity.
Ouroboros fixes the human, not the machine.
What if your AI agent got better just by talking to you?
Introducing OpenClaw-RL — a fully async RL framework that turns your everyday conversations into training signals. Your agent learns your habits, your workflows, your preferences. Privately. Continuously. #Clawdbot #openclaw
🔑 Two learning modes:
• Binary RL — likes/dislikes become rewards
• On-Policy Distillation — your textual feedback becomes token-level guidance
Self-hosted. Zero API keys. Your data never leaves your machine.
👉 https://t.co/ry18qekutm