@yuzu_4ever Start only reading journal papers in ML. TMLR & JMLR are (usually) pleasant to read and very informative due to the quality standards upheld. Still not perfect but remember that a preprint has NOT gone through peer review (altho tbf some of the best work never went to a conf!)
@Keleesssss Yeah super true. I have started using Mathlib + Lean4 for a lot of longer proofs but both Fable + Sol have pretty terrible looking logic to me in my domain. Writing by hand for me is prob faster (bulk of time is reading docs) since LLMs do a lot of trivial/obvious proofs.
@max_spero_@xeophon There is a happy medium (for now) imo. Example: AI plots look terrible but I don't want to read all of the docs for strange edge cases. So if I can write the first ~30%, have a model write the next ~40-60% and then add the final touches myself this is perfect for me. Case by case
@gaunernst@henrylhtsang Even better is when it recompiles from scratch everytime so you have to give it a dummy location in the pyproject.toml for different cu versions for the cache. who knows which one triggers recompile...
@eigenron Finding this out the hard way... Why did no one tell me I would have to look at triple integrals I thought we left that shit behind man. Bring back SDEs please
New blog post: continuous diffusion for language is back!
This research direction receded into the background for a while, but as of this year, it is once again a hot topic. I wrote down a historical perspective and some thoughts on the recent revival.
https://t.co/wNafED6aDB
@cloneofsimo@JulienBlanchon Yes but more conditional infilling like mask git. Like mdlm/bd3lms text diffusion is down to ddpm ~Cat(z), only thing that changes is whether you are in simplex or token space w.r.t pi for the parameterization for the ones in continuous diffusion (~N(z ))
@purshow04 I would say, initially, that a lot of this is valid but the argument of copying images to DDP ranks in specific might not apply. I am assuming you refer to NaViT or newer implementations, but variable resolution/sequence handling has been possible since FA2 with cu_seqlens.
@papayathreesome@ben_golub Я думаию что это меняетсо в других странах. Другие комнани пробайют купить месты па Гоогле. Е-коммерсе всегда также роботал)