@zwcolin LLM always offers some lofty-sounding but practically worthless opinions. It seems as if what we are writing is not a AI paper but a legal document.
I completely agree. The orthogonal Gaussian of LeJEPA is only useful for the same mode (Text/RGB), and it seems that no extraordinary effects have been observed. It is a very ordinary training technique. It is even not necessary. I have verified it in many experiments and even independently attempted to develop similar techniques, but it is extremely, extremely, and extremely unimportant.
I don’t quite understand the motivation for forcing the text embeddings toward an isotropic Gaussian. Language naturally occupies an anisotropic semantic space.
Vision pretraining has many possible downstream readouts, e.g., segmentation, robotics, captioning, localization, etc. Therefore, enforcing an isotropic gaussian distribution makes sense here to best prime the pretrained representations for all possible downstream tasks.
Text captions, in contrast, are already compressed semantic descriptions. Their role here is mostly to guide vision pretraining, i.e., provide a stable training signal for the vision encoder, not to become a general-purpose text representation.
I wish there was an ablation experiment where SIGReg is removed from the text side while kept on the image side. Or perhaps use VisReg (@HaiyuWu1 , @randall_balestr ) on the text side since it gives more flexibility than directly forcing full isotropic-Gaussian structure.
VLM, JEPA and Robots are not very similar to visual tasks, so they are not suitable to be widely used models for visual researchers. We are still expecting other new models.
Overall, the community remains focused on image and video generation, but it is also steadily shifting towards VLMs, multimodal applications, and perception for robotics.
Regardless, exciting times ahead for vision research (provided you have GPUs, the compute report says)!
Yes, even in the Codex era, if you want to create a model that can be popular for 5 or 10 years, you still need to accumulate knowledge and think deeply. The work of some researchers is like a colorful plastic bag.
If you want to ride the wave, you have to see it before it arrives.
The most satisfying moment at #CVPR2026 was when people asked me how to come up with the visionary ideas behind my talk.
Most of those projects were started over a year ago. The best validation is seeing the field catch up.
@wildmindai One of the advantages of Computer Vision over NLP is that its scope is vast (2000+), with countless topics. Each of its topics is a very challenging goal and can leave a strong impression on people through the visual aspect. Just as show in this work.
Why can't @Alibaba_Wan be open-sourced? There are too many mysterious techniques that don't fit well with Wan2.7. If it doesn't become open-source, Wan will also be abandoned by the mainstream technology enhancement solutions.
The open source community has been delivering on LTX 2.3 LoRAs.
Fine tuning LTX 2.3 unlocks control that you can't get from closed models.
Here are 7 LoRAs that can save your footage.
There are too many incredible LoRAs to cover, so comment below your favorite that needs a showcase!
Links to all the workflows below 👇
@SD_Tutorial NVIDIA's approach is similar to one we disclosed two months ago. However, they use DiT to decode in <1 second, while we use CNN to decode in 0.02 seconds.
They didn't notice our work:
Project Page: https://t.co/10K29fIKUD
@wildmindai NVIDIA's approach is similar to one we disclosed two months ago. However, they use DiT to decode in <1 second, while we use CNN to decode in 0.02 seconds.
They didn't notice our work:
Project Page: https://t.co/10K29fIKUD
NVIDIA's approach is the same as the one we disclosed two months ago. However, they use DiT to decode in <1 second, while we use CNN to decode in 0.02 seconds.
They didn't notice our work:
Project Page: https://t.co/10K29fIKUD
Awesome. NVIDIA dropped PiD - fast high-res latent decoding via pixel diffusion!
- replace VAE
- 4/8x upsampling
- 2k decoding in <1s on RTX 5090
- works with FLUX.1/SD3/Z
- rapid generation previews
sharper details, much lower hardware lag compared to standard methods.
https://t.co/60Pkqze0gR
I are very proud that my advisor / the Board of Governors and distinguished Prof. Dimitris Metaxas (#IEEE Fellow, #AIMBE Fellow, #MICCAI Fellow, @RutgersU ) will serve as the #CVPR General Chair to help achieve success for #CVPR2026.