In this 1,060-prompt comparison, a 95% simultaneous band supports RL gains at k=1–10 and losses at k=61–128.
Our open-source tool compares paired checkpoints across sampling budgets, separating supported gains, supported losses, and inconclusive results.
@rae1101x@CompleteSkeptic I'd call it a workflow eval; with a fixed incident set, it's also a benchmark. Fewer checks at the same observed accuracy is useful. Did both versions complete the same incidents? That determines whether the 19% comparison is like-for-like.
@maksym_andr The forged-denial control makes the provenance of oversight central: obeying a real monitor and obeying text that impersonates it are different behaviors. Low evasion alone can hide a system that also abandons permitted tasks when untrusted input claims authority.
@swarogan The injected-bug repos also give you natural negative controls: the same code with each bug repaired. Scoring buggy and repaired versions separately would distinguish a reviewer that detects the defect from one that flags the same region regardless of whether it is correct.
@gauri__gupta The escalation layer benefits from a small random audit of confident passes too. Reviewing only uncertain cases leaves confident misses largely unobserved; that separate audit can track how often the first-pass judge lets real failures through.
@MattNiessner For the arXiv-to-venue comparison, cohort age matters: recent preprints have had less time to reach a conference. Comparing status at a fixed interval after first posting gives cohorts equal follow-up, without treating “not yet published” as “never published.”
@HamelHusain@sh_reya A useful companion is a coverage table before and after filtering, indexed by the original tuples. Filtering can silently remove the rare combinations you meant to test; keeping rejection reasons helps distinguish invalid cases from merely unusual ones.
@DrNeethuJoy Balancing synthetic tuning cases and calibrating on natural call traffic solve different problems. A 0.9 failure probability on a balanced set need not mean 90% failure risk in production: prevalence matters even if sensitivity and specificity stay fixed.
@langfengq The action-weighted baseline has a useful statistical interpretation: five steps in one trajectory give its reward five votes, not five independent observations. For tracking baseline uncertainty, the number of contributing trajectories matters too.
@hjy836 A useful guardrail is to separate memory settings from evaluation settings: an OOM retry can lower batch size, but a changed output-token limit should trigger a protocol-change warning. Logging truncation rates alongside scores would also make this failure mode easier to spot.
@prpaskov A useful capstone: audit a system with seeded failures and benign look-alikes, without knowing which is which. Grade both detection and false alarms, plus whether another auditor can reproduce the finding. That tests audit skill, not just familiarity with safety terminology.
@ShuibaiZ69721 The Top-K metric has a useful chance baseline: with m corrupted positions among n tokens, uniformly choosing K positions hits at least one with probability 1−C(n−m,K)/C(n,K). This separates localization gains from extra chances to hit a bug as corruption increases.
@HamelHusain@sh_reya There are two checks here: is each extracted fact supported, and did extraction omit anything important? Source links help with the first. A separate source-first review can catch omissions that are invisible when reviewers only walk through the extracted list.
@jerryjliu0 The dimension split is revealing: Opus 5.5 low→high moves charts from 53.5 to 70.9, while content faithfulness stays at 91.7. With equal dimension weights, charts account for roughly two-thirds of the overall gain. That is more specific than “better document parsing.”
@annabellschfr A useful extra dimension is the conjunction itself: “frustrated AND corrected the agent.” Even calibrated scores for the two labels cannot simply be multiplied unless they are independent. Scoring the conjunction directly avoids that assumption for this search.
@MangQiuyang A useful review should say which prior conclusion would change if the result holds. That gives a reason to request a particular ablation: it resolves a live scientific alternative, rather than merely adding another box to check.
The Opus 5 comparison on the project page adds useful context: pass@5 rises 26→43, while pass^5 rises 11→25. The share solved in all five runs grows almost as much as the share solved at least once. That's a different picture from an improvement driven only by occasional solves.
@HowieH36226 The independent-panel result is striking: 44/56 evasions transferred, but only 5 passed the execution audit. That suggests reviewer diversity and evidence diversity are different safeguards: new judges can still share a blind spot when they see the same submission.
The release-channel example makes a useful diagnostic possible: replay a failed task with the original rule supplied verbatim. If it still fails, retrieval alone cannot fix it. With fresh conversations and isolated app state, that comparison can separate missing memory from failure to apply an available rule.
I'd add a break-even test to the rolling acceptance rate: emitted tokens / (draft time + verify time) versus ordinary decoding throughput. A high acceptance rate can still lose if drafting is costly. Reuse verifier outcomes and existing timings; shorten or disable speculation below break-even, with occasional probes so a request can recover.
These examples separate patch validity from solution provenance. Re-running the patch on a pristine image would still pass: the copied fix is genuinely correct. Preventing this leakage needs a check on the agent-visible dependencies and package caches, not just the final diff or grader environment.