🔥 Thrilled to share that #SparkVSR was accepted to #ECCV2026!
But before that, our GitHub repo already gathered more than 600+ stars - greatly appreciate the support from the community 👏
As a fully open-source project, we're glad to see that our model was directly supported in-one-click by AI workflow platforms like #RunningHub, https://t.co/6bLYwuxEc1, and Runpod, enabling millions of AI developers and designers to customize their own AI pipelines.
We're also super grateful for the media coverage by https://t.co/lDwIWBTTcM, Hackernews, #Tencent News, to name a few.
We're continuing the R&D of this project and, in the future, will release much lighter versions to support consumer GPUs or mobile devices. Please stay tuned for our latest updates!
☕ Links
1. SparkVSR github: https://t.co/9qVa3rzdla
2. CNAPS AI: https://t.co/jj1fxV2vRW
3. Runninghub: https://t.co/VLpdqThgG3
4. AI Film: https://t.co/FtgGNPXDBW
5. Tencent: https://t.co/iYLoHTWVkA
🎉 Our paper MJEPA has been accepted at #ECCV2026 !!
Huge thanks to my awesome collaborators @AdrienBardes, @michaelrabbat , Sumit Chopra @mattmucklm and Nicolas Ballas
Paper: https://t.co/PmcODf4qld
Code + checkpoints coming soon
See you at ECCV!
Interested in neural SLAM or online semantic scene understanding?
Consider submitting to our #ECCV2026 NeuSLAM workshop!
Accepted or rejected, your recent submission can gain more visibility through our proceedings or nectar track!
More info: https://t.co/fPHDyXpdAV
#ECCV2026 paper: A scalar per patch from pre-trained ViTs enables fast moving navigation in the real world
966 *REAL* nav episodes by S. Janny with Dino-v3, Dino-v2, DUNE, VC1, AM-RADIO encoders show that patch features can be bottlenecked to 1 value ➡️ affordances.
1/8
Thrilled to share that DEER-3D has been accepted to #ECCV2026! ✨
DEER-3D explores whether learning from grounding failures can be more effective than simply scaling 3D training data. We introduce an error-driven refinement loop that identifies predicate-level grounding errors, generates targeted 3D counterfactuals through minimal scene edits, and iteratively improves models with the resulting supervision. Across multiple 3D grounding and scene understanding benchmarks, DEER-3D consistently improves performance.
Updated paper and code coming soon. 🚀
👇🧵
Excited to share our IROS 2026 paper, FLUX: the first flow-based unified policy for cross-embodiment navigation! 🚀
- To overcome the high inference latency of standard diffusion models, FLUX linearizes probability flow into straight-line trajectories, boosting inference efficiency by 47% over prior flow methods.
- By leveraging our new DynBench benchmark and a static-to-dynamic curriculum via GRPO-based RL, the policy also masters socially aware and risk-sensitive dynamic avoidance.
- This enables state-of-the-art performance across six tasks and robust, zero-shot sim-to-real transfer across wheeled, quadrupedal, and humanoid robots without any platform-specific fine-tuning! 🤖
Check out our new DynBench benchmark, paper, and open-source code:
🌐 Project: https://t.co/HShzvu0cRX
📄 Paper: https://t.co/xEUZLOdVwt
💻 Code: https://t.co/lMFRgysfvF
#Robotics #IROS2026 #EmbodiedAI #MachineLearning #HumanoidRobot #ReinforcementLearning
🎉Re2Pix accepted at #ECCV2026!
💡Should a world model predict future dynamics and render pixels simultaneously? Re2Pix says no. Forecast in VFM semantic space first 🧠, synthesize pixels second 🎨
Updated Paper and code coming soon. Details👇
Our #ECCV2026@eccvconf paper
Gen2Balance: Generative Balancing for Long-Tailed Video Action Recognition
now on ArXiv w/ code&data&models
https://t.co/1hhlnbDxZn
Can we fill imbalanced action recognition datasets with generative images to outperform long-tail SOTA?
Yes we can
🧵
🎉 REGLUE accepted at #ECCV2026 🎉
🎨 A unified framework jointly modeling VAE latents ➕ global ➕ local VFM semantics for faster, higher-fidelity diffusion image generation.
💨 Matches 1M-step SOTA in just 700k iterations (~30% fewer steps).
More info below👇
🚀🚀🚀 Excited to introduce AFUN, our first step toward an affordance foundation model for functionality understanding in robotics, led by my PhD student @Zhaoning_Eric_W
Affordance understanding bridges visual perception and physical action, serving as an explainable interface for robot manipulation in open and unstructured real-world environments. Yet, the challenges of building an affordance foundation model are (i) limited datasets to reflect task/environment diversity in the real world, (ii) precise task-conditioned mask understanding indicating where to interact, and (iii) actionable motion to indicate how to interact for robot manipulation. To tackle these, we (i) build a large-scale standardized data pipeline for "affordance" extraction, (ii) leverage a unified model that connects VLMs and SAM3 through MetaQuery, and (iii) use a Bézier Curve representation for motion prediction. The output can be directly deployed on a real robot for execution with zero fine-tuning.
For affordance segmentation, AFUN outperforms all baselines by a large margin across 8 test sets from 4 benchmarks; for contact-point prediction, it predicts substantially more accurate points, with a 12.7–61.3% hit-rate gain over the best baseline; and for 3D motion, it achieves the best performance on all three test sets.
Code & Weights are public. You're welcome to try it out!
Paper: https://t.co/T4FQxEGrMu
Project page: https://t.co/miPx96NgoR
Code: https://t.co/DPDmN17d6N
#Robotics #Affordance #EmbodiedAI #Manipulation #AI
SpectralSplats is accepted at #ECCV2026! 🎉 Tracking 3DGS across frames is harder than it looks — appearance losses break the moment the pose drifts. Spectral moments keep it robust.
@eccvconf
A4D is accepted at #ECCV2026! 🎉 How do you think noisy images are distributed in CLIP space? Surprisingly, they cluster rather than scatter. We leverage this property to detect adversarial attacks simply through the CLIP latent space. Training-free, self-explanatory!🚀
@eccvconf
Thrilled to announce that our paper E3VS-BENCH has been accepted to #ECCV2026!
I'm deeply grateful to all the co-authors and collaborators who contributed to this project.
See you in Malmö!
#ECCV#ECCV2026
Volt is accepted to #ECCV2026 🎉 Huge shoutout to my coauthors @adr_kruse, Tristan Höfer, @dcdegeus, and @BastianLeibe.
The arXiv version uses a NeurIPS/ICLR-style template since we find the ECCV template ugly :)
😺 "The Prism Hypothesis" (UAE)😺 has been accepted to #ECCV2026! 🎉🎉🎉]
🇸🇪See you in Malmö! 📷
💻 Code: https://t.co/vaw74sCkic
📓 Paper: https://t.co/EEx9gL19l3
✨ The core components for image generation and image understanding both lie on the same low-frequency space.
🧠 Paved way for encoder-free design of our Native Unified MLLM (SenseNova-U1)!
🤯 The Prism Hypothesis "v2": Pixel Space Diffusion Made Easy!(https://t.co/LusNSn6OGM)
Our One4D has been accepted to ECCV 2026 @eccvconf. One4D jointly predicts dynamic spatial geometry and temporal video representations in a unified framework, supporting single-image-, sparse-view-, and video-conditioned reconstruction and generation. #ECCV2026
Project page: https://t.co/SUyv1zqs2M
Code page: https://t.co/grntxtlPOr
arXiv paper: https://t.co/D4lIOLtZd2