@ChinmayKak@TianfuF There’s no correct explanation so far I think. Most principled interpretation is some sort of MWU, which makes some sense. But after all this is not something as obviously principled as IS.
@solidSF@yifanzhang_ For a finite game like go, one of the player, black or white, should have guaranteed winning strategy. If you claim to solve go for 7 or 9, can you certify which of the black/white has the guaranteed winning strategy?
@deliprao@harshagundal Well, there is no way you can get true confidence anyway. If the objective is to minimize perplexity, and if the model do really minimize it, then it will return the true class probability and therefore the true ‘confidence’… so i wouldn’t say LLM confidence is nonsense.
@natolambert Have you tried use this approach to scale your RL to use Riemann conjecture? I bet it would keep being all wrong and no signals will ever be detected..