Harness evolution brings a 0.8B model to 100% on ALFWorld. But how much is generalizable improvement vs. overfitting or test-time scaling?
Our method, Harness Delta Attribution, analyzes gains and attributes them to these categories. On 4 benchmarks, lots of the gain is overfitting or due to higher compute usage. Future evaluation should study this more!
Post: https://t.co/QnzjknULFH