Finally, with our findings, we explain why we are able to eliminate the frequency bias in a toy example.
Here taking the distribution of W as a normal distribution with big variance (leading to almost constant rho):
That is, the distribution of the parameter W controls the frequency bias of NN a early times. This also shows in experiments! We tried it on 2 different distributions:
Fortunately, it simplifies if the architecture is chosen as shown in the image, where the W parameter is the same for the cosine and sine. And we arrive at the following result:
Under assumptions, we rigorously describe the error dynamics of 2-layer NN frequencies, focusing on 5 aspects: dimension, Fourier transform of activation g and its derivative g′, and initial parameter distributions. Generally, the resulting dynamics are complex to analyze.
Did you know neural networks have a frequency bias when they learn? Namely, they learn low frequencies before high ones.
We explain this phenomenon and offer tools to eliminate/control it.
➡️ Learn more: https://t.co/Cz9AVM5mf8 🔥
Thread 🧵 (1/7)
@fcosahli@MirceaSci