I don't imagine this is news to people with a background in simulated annealing, and I would love to get some input on what annealing schemes are actually used in practice nowadays! 7/n
I have been long wanting to compare diffusion models to MCMC empirically.
Running Langevin using the annealing scheme from diffusion models results in good samples even for hard targets. Details at https://t.co/lG5jimxCud
1/n
We can quickly create an annealing scheme that replicates that behaviour with the ansatz p_t = p(alpha_t x)^beta_t. This already works much better than the scheme employed before and is implementable.
6/n
I am quite annealing schemes of this form already have to already. Would be very interested on any input on what annealing schemes people actually use nowadays in practice. 6/6
However, one can try to approximate such an annealing scheme by choosing p_t = p(alpha * x). A very short first attempt already improves upon the above annealing scheme! 6/n
Intuitively, the modes for very smoothed out versions of p_t are all at 0. Then the modes move to their target destination, transporting the samples along with them. Unfortunately, for more complicated distributions, this annealing scheme cannot be computed. 5/6
For a Gaussian mixture, one can calculate the solution to the forward SDE in diffusion models/score-based generative models in closed form. Therefore, we can use this path as smoothed versions and plug it into a Langevin SDE. Turns out, this works nearly perfect! 4/6
@david_picard That is not what is discussed in https://t.co/aROGPwGvfI, but it shows that SGMs do indeed memorize the training data if you train them for long enough. The implicit regularization by early stopping of the training seems to be crucial.
@david_picard One quick experiment to see it overfitting in action: Running it once will trigger the SGM to generalize and generate new points on the sphere (the toy distribution is the uniform distribution on the sphere).
When you increase N_epochs, it will memorize the training data.
The Jupyter notebook here contains an introduction to SGMs using JAX: https://t.co/FHFuWVj0Sv
It sets a focus on understanding some theoretical aspects as well as their generalization capabilities.
@JamesTThorn@david_picard This is an very nice paper employing an elegant trick to get convergence results. However, the results only hold in the case where p_data > 0 everywhere. To study memorization, one needs to choose p_data as the empirical measure, which is only supported on finitely many points.