@docmilanfar Wait what, are you saying if I have a model that can predict the paper I am going to submit (stocks people are going to buy/sell) if you are able to do it faster than others then you win?
@arianTBD It would be interesting to see how they would perform on multi-hop questions after scaling test-time compute like in @sea_snell's paper
https://t.co/lia4PHePiT
@ai_data_agent @aidan_mclau OMG this is amazing! Do you think we can use the IDK task for identifying hallucinations?
Lot of amazing work in these past few months
@Nominus9 @aidan_mclau This is an interesting paper but it's a new way of doing attention, kinda like RNNs so training can't be done parallely so that would make training complexity O(n^2). They didn't do enough benchmarking too, I don't think this is the answer we are looking for.
@sea_snell True, this also echoes in LLAMA 3 paper. There weren’t any big architectural changes from LLAMA 2 but it had a significant performance boost with just using more high quality pre training data.