@cloneofsimo did your cat take into account that even if the DFT weights increase with number of pixels, the mean square error also should divide by the number of pixels, since it's a spatial mean? perhaps preventing mse from blowing up?
@ostrisai interesting! what is the convergence speedup vs say Adam? also I wonder about the connection with running average since negative + positive numbers will average to 0
@KBlueleaf in parallel settings like FSDP, doesn't the lower ddr memory generation in epyc 7003 series & slower cpu memory controller bottlneck your communication between GPUs?
@cloneofsimo I think you are right, theoretically. You can't ignore the fact that CoT definitely works really well, but I think the reason it works might currently be not so well understood and over-anthropomorphised. it might just come down to using more flops or more functional complexity
@Mascobot@a16z Have you tried enabling peer to peer PCIe communication between 4090s?
My understanding is sometimes the bottleneck is not in the actual PCIe bus but the memory controller of the CPU, so going through RAM is the bottleneck and you could get full PCIe utilisation by using p2p.
@jon_barron all images and videos are 3D data, since are projections of 3D a world. if we can insert the right inductive biases into image models maybe we implicitly learn good 3D foundation models with no 3D annotation.
arguably current models are already doing so to some extent.
after running the road runner scenario through a monocular depth estimation model (that probably has much less accuracy than the real FSD system would have).