FoPE demonstrates strong performance in several scenarios, including training from scratch, continual training from a well-trained RoPE-based model, and integration with well-known extrapolation methods (e.g., YARN).
π The Source Code of FoPE is now released:
https://t.co/L036gb8Mtn
It is interesting that not only does attention influence length generalization, but linear layers and activation functions also play a role!
how did this go under the radar?
less noisy long context generalization is very nice. note that it gets better at 1024 token sequences despite being trained on 512 tokens
(note that ALiBi doesn't extrapolate meaningfully in practice bc of how it decays)
@YouJiacheng What's your evidence for your understanding is how FoPE or RoPE works?
Let's save our times to read more papers or deepen our basic knowledge. thanks
@YouJiacheng We appreciate the works you mentioned. But we believe necessary theoretical modeling is also important (even less important than empirical result).
We will continuously work to prove or falsify the theory. But we will not reject making potentially valuable assumptions.
@YouJiacheng Thank you for your comments, but it is so regretful that it is totally not the way how FoPE or RoPE works.
I sincerely suggest you may have a more careful understanding of these methods and their background knowledge before comments.
@musical_chemist Thanks for your interest in our work. We will open source our code as soon as possible in Github: https://t.co/L036gb8Mtn
More details can also be found in Hugging Face Daily Paper: https://t.co/f7hPYWXNJU
@kalomaze Thanks for your interest in our work. We will open source our code as soon as possible, once the main contributors of FoPE have finished their exam period.
Github: https://t.co/L036gb8Mtn
HuggingFace Daily Paper: https://t.co/f7hPYWXNJU