@natolambert So question: why does a model tries to hack two orgs instead of answering the question? Which part of the training leads it to do that. The valid examples ive seen are when the questions are impossible.
@Suhail I'm just arguing that you don't need logits to do distillation & have a big perf impact downstream. FWIW I agree with the exaggerated part of your statement.
@yongfook The value of a phone for a billionaire and a normal person is vastly different yet cost is the same. Cost # value add it’s more about competitive structures
@jonchu The tough bit is the data behind the environment/api and the interplay with api logic. Infinite complexity there combined with data staleness as soon as your api change. Similar challenges as testing.
@khoomeik “try coming up with an order that’s more authentic”
“Here is my new order”
Reward=1
Hope it generalises.
You’ve solved RL’s biggest flaw with feedback aware sampling.
@dwarkesh_sp@gwern I think the missing value add of podcaster is the amplification of (potentially already available) perspective. I listen to you because I trust you to amplify perspectives that I find interesting. My goal is not to uncover information not available anywhere else.
@hardmaru@SakanaAILabs How about applying to the automated production of reproducibility reports of existing papers? This should be easier than producing entirely new papers, and be quite useful in that this is a type of work that we probably don't do enough as a community.