๐งต Understanding GRPO in DeepSeek R1: for Efficient Training ๐
Traditional RLHF uses Proximal Policy Optimization (PPO), which is unstable and inefficient. GRPO simplifies the process by removing the KL penalty and directly optimizing against the reward model.
Even exams for patwari/clerk/startups have eligibility criteria, and they get fired for underperformance.
I always believe patriotism = terrorism, or one man's freedom fighter is another man's terrorist.
Maybe I'll write a blog on this West led "StupidCracy," later. ๐
Nepal will repeat itself.
They hammered two nails into their own coffin: kidnapping Sonam, then spreading love to protesters on the 20th.
China deployed the army at Tiananmen Square.
No dictator ever survived their own tyranny.
Even I/we would've misused that much power, money, and fame.
Decentralization and more autonomous pillars are the solution.
Apart from legislative, judiciary, executive, and media, where hijacking by one is too hard.
@dwarkesh_sp Yeah!! I don'watch anyone else pod, and I use him @lexfridman as a benchmark,i totally don't understand: in present value created by him vs his past- is he associated with mit or not, rubbish!! ๐ I love lex and we r here in your hard times....โค๏ธโค๏ธ
An exciting milestone for AI in science: Our C2S-Scale 27B foundation model, built with @Yale and based on Gemma, generated a novel hypothesis about cancer cellular behavior, which scientists experimentally validated in living cells.ย
With more preclinical and clinical tests, this discovery may reveal a promising new pathway for developing therapies to fight cancer.