15/15 Verified fixes are the currency.
Keep the receipt: what changed, what passed, and what still needs a person.
Plugin: https://t.co/W88kcn6Eqd
CLI: https://t.co/keH1eUJlGg
Session: https://t.co/8xUKCtHkHZ
14/15 Put the fine print in the plan:
• CLI and SDK are public; access and hosting vary.
• Keep unrelated credentials out of scans.
• Reports may contain code and vulnerability details.
• Deployed systems still need their own testing.
13/15 Our recommended pilot:
1. One critical repository you own, with a named reviewer.
2. A focused fix: reproduce, patch, test.
3. An advisory scan: measure signal quality, cost, and delay.
4. An enforced gate: expand when the workflow earns it.
1/15 The scanner found it. Now show the fix works.
Our takeaways from OpenAI's Daybreak Live: validate the finding, test the patch, and put the decision in human hands.
A 15-slide guide to the evidence between a finding and a reviewed fix.
12/15 Pick the model. Pick the effort. Watch the dollars.
An estimated scan budget is not a hard spending cap. Requests already in progress can finish above it.
Choose settings against the coverage and turnaround time your team needs.
11/15 Then it waited for a human.
In the demonstrated workflow, the agent opened the pull request and notified the engineer in Slack.
The engineer supplies context, sets the threshold, and reviews the diff. The handoff includes the evidence.
10/15 Repository rules make the gate real.
A finding above the configured severity threshold fails the scan check. To block a merge, the repository must REQUIRE that check.
Annotations alone do not enforce the gate.
9/15 From one repository to every pull request.
The demo showed 1 repository in the desktop app and 2 selected from 32 listed. Pull requests can be checked through configured CI.
Evaluate signal quality, runtime, and cost before enforcing a gate.
8/15 An unfinished scan is unfinished work.
Exit code 2 means the scan hit an error or left coverage incomplete.
Review partial or unknown coverage. Check deferred areas before relying on the result. A finished command is not always a complete review.
7/15 The patch ships with its own receipt.
1. What changed: a focused patch.
2. What was tested: before-and-after results.
3. What remains: coverage limits and proof gaps.
Passing tests are evidence. The engineer still makes the merge decision.
6/15 Evidence before urgency.
Look for an attack path, a source-to-sink trace, and a severity assessment of reachable impact.
The Ladybird demo involved about 28,000 files, per the session notes. The size matters less than evidence a reviewer can inspect.
5/15 Give the scanner the context it needs.
SECURITY.md should explain:
• Scope: what belongs in this review?
• Existing controls: what already limits risk?
• Business logic: which rules must the system uphold?
Useful findings start with clear boundaries.
4/15 The work between the work:
Context → Find → Validate → Fix → Verify → Human review.
The reviewer decides what ships. That decision feeds context into the next cycle.
Automation moves the work. A person owns the decision.
3/15 Two questions deserve separate evidence:
Is it real? Trace a reachable attack path.
Is it fixed? Retest the issue and record the gaps.
False positives waste review time. Regressions create more work.
2/15 OpenAI reported 30M+ commits scanned across 30,000+ codebases, with 500,000+ findings automatically determined to be fixed.
Those June 22, 2026 figures do not establish how many AI-written patches reached production. Read the definition with the number.
6x growth YoY for @Airbnb stopped chasing tourists & started selling to locals. Paris is converting local residents into bookers, a segment rival platforms have written off as unprofitable. @Roanoke_Region gig economy level programming. Let’s build!
https://t.co/ZpCf6Akzit
Business owners in the Roanoke Valley, Blacksburg & New River Valley: I’m building an AI readiness curriculum for small and midsize businesses.
Which topic would help you most? Vote, then reply with one task you’d like AI to help with.
Turn the dial before you switch the model.
Claude Opus 5.5 has five effort levels that control how much thinking it puts into a task.
Results falling short? Give it a way to check its work, then try more effort before switching to a more powerful model.