@Laz4rz You should probably also know the result that tells you why it works in the first place (minimizing MSE naively would cause you to predict the mean each time and lead to a blurry image)
@yule_gan Perhaps it’s because pretraining creates a set of experts sufficiently close to the model in weight space such that a random perturbation can move the model onto the an expert? This paper has more details: https://t.co/6ba84paUPB
@avivbick@rshia_afz@CevherLIONS@ericxing@_albertgu Do you think this advantage over SSMs will scale up well? It seems like a sufficiently trained SSM will be able to predict a delta t vector that approaches 0 on the information it wants to preserve.
Current video autoencoders waste enormous amounts of capacity on uninteresting information; a video zooming in on a rock contains much less information than a lion running on a Savannah. Here's a way to fix it.
https://t.co/0FjIKECHm8
@Memetic_Theory Google has a TRC program that gives you 64 v6e TPUs (a bit over 32 h100s of compute) for a month. They’re quite generous, and you should be able to get it if you apply.
@gabriberton This is a fairly common thing with a lot of LLMs, especially Qwen. RL on itself narrows the model’s distribution which reduces the probability of an error up to an extent.
@danpacary There’s no need to do LoRA by the way. If you’re not computing gradients, each set of weights can be represented as a scalar seed, so LoRA just reduces expressivity without any memory benefit.
@nabla_theta Even in a potential future highly competitive market, there will probably be some AI companies that are more trusted than others because some companies have a large number of highly opinionated researchers that would leave if the company did something terrible.