You shouldn't have to click through every button on every screen in your app to find that broken edge case.
We find every bug across security, ui, payments, auth, or api before it finds your users.
Run your first scan now @ breken[.]ai
Your users hit bugs every day, and almost none of them tell you. AI agents use your API every day, face a bug, and move on.
▚▚ Scout catches both. It asks frustrated users what went wrong, lets agents report failures in a structured way, verifies each one in a sandbox, and hands you the fix.
This is the beginning of self-healing software. Software that heals itself based on what its users actually tried to do.
https://t.co/CiEoyzzLBe
We pointed our bug-finding engine at some of our favourite open-source repositories.
We got merged into Docker, Supabase, Langflow, Hugging Face, and about 40 others that are changing how we interact with software.
The truth is none of our codebases are perfect, but Breken can get you closer.
Have a look @ https://t.co/nyIj0EqRNZ
@i_mika_el graded against the code after the fact. the tag is the model's claim about how it got the answer; we then check the answer against the actual source, and a wrong inferred demotes the binder that produced it. self reported alone would be vibes.
@MichLieben a 99 health score that doesn't predict spam placement is the same failure as a stale index, the number looks fine but measures the wrong thing. what helps is reporting how far the view has drifted instead of asserting it's current.
@tsapeta the annotated Swift class is doing upfront what our resolver does after the fact, giving the agent ground truth instead of a guess. can point our analysis at your repo and send back what it finds, if that's useful.
@bil0090 16 hours unsupervised and it found bounties, that's the easy half. the hard half is knowing if a completed bounty actually passed the repo's own checks or just looked done, a different model reviewing the diff catches that gap.
@SammyTourani@Wealthsimple 200 PRs across 63 repos in 4 months is real throughput. eval harness grading the reviewer is the piece most teams skip, the gate that stops a model marking its own homework. happy to run ours over yours and send what it finds.
@sanketsaurav we adjudicated 58 findings from another reviewer's run and 56 were marked P1, severity carrying zero information. verification only works if the tool can say some things checked out and this one didn't. happy to run it over your repo and send back what it finds.
@tomjohndesign running 10 agents merging nonstop is where risk actually shows up, not PR one. make two-of-three verification gates a refusal, not a merge, so failures don't compound. can run it across yours and send back what it finds, free.
@BkashJosi reorganizing is easy, proving it's safe isn't. rescuing 6 files back introduced 6 compiler errors absent when all 909 were deleted together, a rescued file still imports dead neighbours. can run the same check on yours if useful.
@emanueledpt agents export by habit, not need. one scan found 906 dead exports, but only 31 held up once it checked usage three lines below the declaration. glad to point ours at your repo and send back the real list.
@DanKornas the gap that breaks this setup isn't planning to code, it's code to trust. two out of three verification gates passing still isn't a merge in our system, it's a refusal. happy to run ours over your repo and send back what it finds.
@matthewmillerai One-shotting a 6-month bug means little if nobody re-ran the repro after the patch to confirm it's actually dead. The models that gave up may have been more honest than the one that quietly shipped a fix that didn't hold.
@Av1dlive shared context across models breaks fast once one of them assumes the wrong language. a TS parser once got pointed at a Go repo and returned garbage symbols while tsc stayed green, worth checking that edge before trusting the brain.
@titouangalopin a preview env plus a self written test plan still means the agent grades its own homework. the gap is who reviews the plan itself, not just runs it. we'll run our three-gate check on your repo for free and send back what it catches.
@HashgraphOnline scanning before execution catches syntax, not intent. the harder problem is stopping self-review bias, we make it a compile-time guard so the fix agent literally cannot read its own score, not just a policy nobody enforces.