Does an imperfect verifier break reinforcement learning with verifiable rewards (RLVR)? Turns out it doesn’t!
Why does this matter? As the world moves into reinforcement learning in semi-verifiable domains, perfect verifiers don’t exist.
We added controlled and LLM-based noise to RLVR reward signals and found that up to 30% noise barely hurts training; performance stays within 4pp of the clean baseline.
This research has already impacted how we build reinforcement learning environments at @joinHandshake. For a major benchmark we are launching tomorrow, we hill-climbed the verifier to 88% accuracy—above the 85% human inter-rater agreement—knowing from this research that this is good enough.
With @andreas_plesner@guzmanhe
News: @joinHandshake acquires @CleanlabAI!
This "ten-year old job marketplace" has quietly become a top human data lab for AI--building an AI research org, acquiring top AI talent, and advancing Cleanlab tech and research to lead data foundations for frontier AI.
1 of 4
.@CleanlabAI has just been acquired by @joinHandshake!
This “recruiting marketplace” has silently grown in just 1 year to be a dominant player in human data for AI.
With Cleanlab's deep roots in research, Handshake is doubling down on building out its AI research org to strengthen the data foundations for frontier AI.
🚀 New from Cleanlab: Expert Guidance
AI agents running multi-step workflows can fail in tiny, trust-breaking ways.
Expert Guidance lets teams fix these behaviors with simple human feedback, instantly.
✈️In one airline workflow: 76% → 90% after only 13 guidance entries.
“A great soul never dies. It brings us together again and again.” — Maya Angelou. You’re always with us Steve, your memory connects and inspires us every day.
Don't let smiles fool you...these 2 wld knock you over for a lose ball, drain a three on th other end, & TELL you all about it! #Love&Bball https://t.co/7yaBeYuCc2