Indian developer reverse engineered the protocol that AirPods use to communicate with Apple devices and implemented those features on Android (and Linux) |
Announcing Together Chat!
Use DeepSeek R1 & other top open source models to do web search, coding, image generation, & image analysis.
Available today for free!
I believe that a misunderstanding of the formula for GRPO is currently spreading, because people don't actually look at formulas in detail. can someone please check my sanity? because surely someone else would have noticed, if I was right.
in the deepseek-r1 paper, square brackets are used to denote what we take the expectation over (i.e. the domain of the expectation), which is of course highly unusual and confusing, as we usually put into square brackets what we take the expectation of. this leads many people to falsely believe that we have the expectation of q ~ P(Q) and {o_i}^G ~ \pi(O|q) times the 1 / G * sum term.
on the off-chance that I am indeed not wrong in my understanding, I believe that it would be beneficial for the sake of clarity if the authors would change the formula to fit the standard convention by turning that which is currently in the square brackets into a subscript of the expectation and putting the other term into square brackets.