Giveaway time! Here’s your chance to win the Osmo Action 6 Standard Combo.
How to enter:
1. Follow @DJIGlobal
2. Like and share this post
3. Bonus Chance: Comment below what you’ll be filming with this camera in 2026.
· Time period: 2026/3/6 - 2026/3/31.
🚀 New paper: https://t.co/ei1nic0fqZ
VideoLMs are bottlenecked by a simple problem: they treat video like a stack of images. That means huge token costs, slow responses, and missed temporal details.
What if we processed video the way codecs do? 🎬
Instead of dense per-frame RGB embeddings, we tokenize motion vectors + residuals and only encode sparse keyframes — turning video redundancy into a powerful inductive bias for efficient temporal reasoning.
🧵👇
🚀Announcing C-RADIOv4: The latest evolution in our agglomerative vision backbone family is here!
We’ve built a unified student model that distills the best capabilities of multiple state-of-the-art teachers into a single, efficient architecture.
DINOv3, SAM3 and SigLIP2 all in one forward pass with better features. Better efficiency, robustness and any resolution.
👐Permissive and open-source!
Read the tech report: https://t.co/nopJ7zxJte
Models:
https://t.co/5Pa0kVpxap
https://t.co/H8T8bn67MX
Github: https://t.co/YxhkCfDKLd
🧵1/N
🚀New from Meta FAIR: today we’re introducing Seamless Interaction, a research project dedicated to modeling interpersonal dynamics.
The project features a family of audiovisual behavioral models, developed in collaboration with Meta’s Codec Avatars lab + Core AI lab, that render speech between two individuals into diverse, expressive full-body gestures and active listening behaviors, allowing the creation of fully embodied avatars in 2D and 3D.
These models have potential to create more natural, interactive virtual agents that can engage in human-like social interactions across a variety of settings.
Learn more: https://t.co/vl2jmfE7wX
1/2) Happy to share the preprint of our workshop paper on using information theory to find class separation in diffusion models
It generalizes previous models of speciation and symmetry breaking to generic class definitions
Did @Google just release a better version of SigLIP?
SigLIP 2 is out on Hugging Face!
A new family of multilingual vision-language encoders that crush it in zero-shot classification, image-text retrieval, and VLM feature extraction.
🧵👇
A lot of my phd I tried to frankenstein together many of these popular tools.
Only to later waste many days debugging from some bug introduced by these tools
Join us at the @CVPR workshop on 'What is Next in Video Understanding?' to hear from our excellent line-up of keynote speakers.
We also have an open call for 1-2 page position papers on the future of video understanding.
https://t.co/QOPSsoDCIS
#CVPR2024
Want to make sense of large embeddings? WizMap it!📍
WizMap is an interactive visualization tool for exploring embeddings. Seamlessly navigate through millions of points, while gaining valuable insights from multi-resolution summaries!!
👉 Try it now: https://t.co/in46KhGrRY
OMG YES!! please do it. I’ve been frustrated with AMD software being so bad and missing out on DL ever since I started, which is over a decade ago. Because their hardware is/was dope.
Seriously considered gambling my phd on doing this myself, though ended up deciding against.
Today we release LLaMA, 4 foundation models ranging from 7B to 65B parameters.
LLaMA-13B outperforms OPT and GPT-3 175B on most benchmarks. LLaMA-65B is competitive with Chinchilla 70B and PaLM 540B.
The weights for all models are open and available at https://t.co/q51f2oPZlE
1/n
What are the best ways to make self-supervised visual representation learning more efficient?
At #NeurIPS2022 tomorrow, we’re presenting new research showing how to evaluate the compute efficiency of popular pre-training methods. https://t.co/mzfoZHPqko