Introducing @QuiverAI, a new AI lab and product company focused on frontier vector design.
We’ve raised an $8.3M seed round led by @a16z, with support from amazing angels and investors.
Our first model, Arrow-1.0, generates SVGs from images and text. It’s available now in public beta at https://t.co/zjAnKlI8pp
🎓 Just wrapped up this semester’s “Diffusion and Flow Models” course at KAIST, covering DDPM/DDIM to SDE/ODE and the connection between diffusion and flow models. Next week we host guest talks by Subham Sahoo and Philipp Henzler.
🌐 https://t.co/fYoEth6iM8
I’m excited to announce that 💫StarVector has been accepted at CVPR 2025! Over a year in the making, StarVector opens a new paradigm for Scalable Vector Graphics (SVG) generation by harnessing multimodal LLMs to generate SVG code that aesthetically mirrors input images and text. With this milestone, we’re also releasing StarVector on Hugging Face! 🥳🚀
🔍 Fascinating discovery: Diffusion models perceive light/color as humans do!
In our recently accepted #CVPR2025@CVPR paper with @Alvaro_kai93 , We've found that during inversion, diffusion models naturally replicate human responses to color visual illusions.
🧵
Introducing REPA! We show that learning high-quality representations in diffusion transformers is crucial for boosting generation performance. With REPA, we speed up SiT training by 17.5x (without CFG) and achieve state-of-the-art FID = 1.42 using CFG with the guidance interval. 🧵[1/7]
https://t.co/MoN9vjZ93U
arXiv -> alphaXiv
Students at Stanford have built alphaXiv, an open discussion forum for arXiv papers. @askalphaxiv
You can post questions and comments directly on top of any arXiv paper by changing arXiv to alphaXiv in any URL!
This is on of the best papers I have recently read, congrats for this excellent work.
"NNs do not have an inherent “simplicity bias”. This property depends on components such as ReLUs, residual connections, and layer normalizations".
This also applies to transformers.
one of the most important things I know about deep learning I learned from this paper: "Pretraining Without Attention"
this what I found so surprising:
these people developed an architecture very different from Transformers called BiGS, spent months and months optimizing it and training different configurations, only to discover that at the same parameter count, a wildly different architecture produces identical performance to transformers
this may imply that as long as there are enough parameters, and things are reasonably well-conditioned (i.e. a decent number of nonlinearities and and connections between the pieces) then it really doesn't matter how you arrange them, i.e. any sufficiently good architecture works just fine
i feel there's something really deep here, and we may be already very close to the upper bound of how well we can approximate a given function given a certain amount of compute. so we should spend more time thinking about other questions, such as what that function should actually look like (what data? which objective function?) and how to make it more efficient
And the winner is...
🏆@hector_laria for the best presentation (2nd year in a row)
🏆 Pau Torras for the best quiz score (2nd year in a row)
🏆 Khanh Nguyen for the best poster
Congratulations!!👏🎉
#CVCpeople#phdlife
@PDillis@quasimondo Related papers are Latent Diffusion Models and Arno's group paper on heat dissipation models, where blur is used instead of noise, and the initial color mass is kept for the final image and just redistributed 🧐 https://t.co/Wxc5mxogmb
@PDillis@quasimondo Generally, domain specific knowledge is sideloaded during diffusion, i.e., controllable score-based models. But hey what do I know, you could find something interesting!
📢📢 The recording of our tutorial on denoising diffusion models is now available on YouTube:
https://t.co/JB7SLuvxuW
Slides and additional info:
https://t.co/5H4zFeovF2
This tutorial was originally presented by @RuiqiGao, @karsten_kreis, and myself at #CVPR2022.