I'm entering the OpenAI DevDay Side Quest sweepstakes for a chance to win an OpenAI prize! #OpenAISideQuest
Checked my ChatGPT profile and one stat floored me: 2.2 BILLION lifetime tokens, with a 4h 43m longest task. Either I'm very productive or the machines are keeping me very busy. 🤖
Solo groomers lose ~2 hours a day to "can I book?" WhatsApp messages, and fixed time slots don't work when every groom takes a different time.
For the #lovablechallenge, I built Groom Room by Asha with @Lovable, a booking app for a solo dog & cat groomer in Bengaluru.
What it does:
• Works out each groom's length from breed, coat, matting and service (30–150 min), and shows only the slots that truly fit
• Vaccine proof upload with expiry checks and a pending hold until the owner verifies
• ₹200 Saturday deposit against no-shows
• Customers reschedule or cancel in one tap, and freed slots go straight to the waitlist
• Owner dashboard with today's timeline, care notes, vaccine checks, rebook reminders and time saved
In the demo week, the owner saves ~190 WhatsApp messages and 4+ hours.
Built with one detailed prompt, then fixes in batches, on Lovable Cloud. Every flow was tested on the live app.
Try it: https://t.co/YVUZTKMosv
Full walkthrough in the video 👇
#buildinpublic #nocode #AI
step 4 is the whole exercise honestly. did something like this on a real PR to js-x-ray recently - my weak-pbkdf2 detection looked fine, 849/849 tests green, and the reviewer still found two real bugs: sha1 needed its own 1.3M OWASP threshold, and my digest check was case-sensitive. both got fixed with regression tests. verify every line applies to your own tests too
the interesting bit is the shared context. reviewing with the same agent means asking "why did you do it this way" gets a real answer instead of a fresh hallucination, because it already knows its own reasoning.
also $700 for one PR - at that price the review chat better be worth it.
@fourweekmba@tryramp 75% opened is not 75% shipped though. the caveat at the end is the whole story - agents inherit the architecture. from the OSS side: repos with brutal CI absorb agent PRs fine, everywhere else the merge queue just gets noisy.
@automater_ai@askalphaxiv this matches what i've been doing with jev - cheap fast decisions for the routine calls, bigger model only when confidence drops. making the threshold visible is the part most people skip.
the confident-and-wrong case is the scarier one. i run jev as a gate over my own agent's tool calls and the pattern i keep seeing: a confident wrong verdict shuts down further reasoning, while an unsure one at least triggers a second look.
the "judged the code or bought the story" framing is exactly right. a confidence score means nothing without the evidence trail next to it.
moving from review to change management makes sense, but review only gates what the test suite lets it. on my recent OSS PRs the maintainer review plus green CI did the real gatekeeping, bot comments were secondary. agentic change management only works if the test suite is what everything defers to.
tried a similar setup for browser automation. jev makes one fast decision per call, my agent supervises with a success check and a step budget, and a stale-decision retry catches it when jev starts repeating itself. the millisecond decisions are the win, the supervision layer is what makes it usable.
the llvm policy is the only honest one. no free pass.
on my first merged OSS PR the reviewer caught two things i was sure about: sha1 needed its own 1.3M OWASP threshold, and my digest check was case-sensitive. i wrote regression tests for both fixes.
AI writing the code doesn't change the requirement. you still have to understand every line.
@0xShoopy the tests-that-repeat-the-code point hits home. on my first merged OSS PR the maintainer's review found two real bugs and i wrote regression tests for both fixes, 849/849 green. that's the gap between tests that catch something and tests that pass forever and prove nothing.
this maps to what i see running jev as the fast layer under my own agent loop.
→ ordinary judgments: jev handles them fine, costs almost nothing
→ anything that needs an actual derivation checked: i never let the cheap layer decide, the supervisor re-verifies against the goal before committing
the interesting number isn't the discount, it's the false-confidence rate. my loop has no-op detection and stale-decision retry precisely because the cheap model will happily keep going while wrong.
the hybrid part is what makes this work. pure LLM review misses deterministic stuff like hardcoded salts and iteration counts, pure static analysis can't reason about intent.
i built a weak-pbkdf2 detection probe for js-x-ray where this exact split showed up: fast checks flag candidates, judgment decides what actually matters.
one of my recent open source contributions just made it into a real release.
i contributed to "skfolio/skfolio", a 2.4k+ star open source python library used for portfolio optimization, quantitative finance, risk management, factor models, backtesting, and model selection.
repo:
https://t.co/WrVhgZWnEF
my contribution was around "MultiPeriodPortfolio", which is used when you evaluate or combine portfolio results across multiple time periods.
the bug was related to "sample_weight".
after operations like append, delete, or replacing a portfolio period, the number of observations could change while "sample_weight" still had the old length.
for example:
4 portfolio observations
but only 2 sample weights
that could leave the portfolio in an inconsistent state and later cause weighted calculations such as mean or risk measures to return invalid results.
my fix added validation before the mutation happens, so incompatible updates are rejected before the portfolio object gets corrupted.
i also added regression tests covering:
append
delete
different-length replacement
same-length replacement
portfolio list assignment
valid weight reassignment
the maintainer later expanded the implementation with inherited sample weights and broader weighted-measure handling.
the PR was merged into the real upstream repo and shipped in "skfolio v1.3.2".
PR #338:
https://t.co/R4OAI5kNhl
this is the kind of open source work i want to keep doing.
real quant library.
real bug.
production code.
tests.
maintainer review.
merged.
released.
review it. harder than a human's, honestly. ai code reads clean on the surface and hides the subtle logic mistakes underneath. i've been on both sides of oss review this month, the reviewer catches what the author can't see, and authorship doesn't change that. trust the tests and the review process, never the author.
the loop is the interesting part, not the fuzzer. fuzz → measure → improve → repeat is exactly how a human researcher works, just automated. most ai security demos stop at generating an answer; this one observes coverage gaps and changes strategy. and the triage pipeline is what makes it usable - raw fuzzer output is hundreds of duplicate crashes that nobody reads.
the irony at the end is real though. you give the agent shell access to compile and run arbitrary build commands so it can find vulnerabilities, and now the security tool itself is the attack surface. disposable environments aren't optional here, they're the whole design.
@shneural the unprompted blender decision is the whole story. same access, same prompt, one model looked at the tools and did nothing with them. we keep benchmarking what models can do when asked. the gap is what they choose to do unasked.
@shubh19 step 4 is the whole game. on my last OSS PR the human reviewer caught two things no test flagged: a case-sensitive digest comparison and a threshold that needed to be per-algorithm. tests were green, the bugs were real. reading the diff line by line isn't paranoia, it's the job.