Huge congrats, this is super exciting!
Really cool to see this direction pushed so far. We explored a related idea back in 2023 with Unleashing Text-to-Image Diffusion Models for Visual Perception (https://t.co/qNZIbYdZuW), where we studied how to leverage pretrained text-to-image diffusion models for downstream perception tasks. Using Stable Diffusion 1.5, that work already showed how far pretrained image generation models could go for perception, including SOTA results on NYUv2 depth estimation, well before follow-up works like Marigold.
Excited to see this line of work evolve at GDM 🚀
This is a thoughtful piece of work on a question that deserves more attention: what makes a latent space good for planning, not just for representation quality? The key idea, temporal straightening, is compelling because it explicitly shapes the geometry of latent trajectories so that optimization in latent space becomes better conditioned. A very clean and insightful perspective on latent planning. Big congratulations to @yingwww_ and the team on this excellent work.
What is a good latent space for world modeling and planning? 🤔
Inspired by the perceptual straightening hypothesis in human vision, we introduce temporal straightening to improve representation learning for latent planning.
📑: https://t.co/CCmcEIJGM6
🚀 Excited to share that our paper “CapNav: Benchmarking Vision Language Models on Capability-conditioned Indoor Navigation” has been accepted to #CVPR2026 (main track)!
📄: https://t.co/ZdSpsMPBJY
💻GitHub: https://t.co/1XkvpxywEY
🤗Hugging Face: https://t.co/LzQ77EykX4
(1/4)
We found that visual foundation encoder can be aligned to serve as tokenizers for latent diffusion models in image generation!
Our new paper introduces a new tokenizer training paradigm that produces a semantically rich latent space, improving diffusion model performance🚀🚀.
I think the problem is that most academic 3D work stops at learning a representation. But in real-world commercial applications, a representation is meaningless on its own — its value only emerges through downstream output formats like language, action, or pixels. Without that grounding, we’re just building beautiful but hollow spaces.
Nobody pays for a latent space. People pay for what they can see, say, or do.
Can data owners & LM developers collaborate to build a strong shared model while each retaining data control?
Introducing FlexOlmo💪, a mixture-of-experts LM enabling:
• Flexible training on your local data without sharing it
• Flexible inference to opt in/out your data anytime
At 37B parameters, FlexOlmo is competitive across 31 tasks.
Having trouble dealing with the excessive token number when processing a video? Check out our paper that is accepted by ICCV 2025 with an average score of 5.5! We tokenize video with tokens grounded in trajectories of all objects rather than fix-sized patches. Trained with a CLIP objective, TrajViT is the first model to consistently outperform standard ViT with all video understanding tasks we tested even with modern Video-LLM QA—while using ~10 × fewer tokens.
paper: https://t.co/aewxhdAgI1
project page: https://t.co/9KbOxwXoSU
Details in the thread.
I need to read it carefully, but now this IMO is likely most deserving of "important papers in LLM RL since R1".
If you try 100 random underpowered tricks and all of them lead to huge gains, but only on certain model class X, the finding is about X, not about the random tricks!
🤯 We cracked RLVR with... Random Rewards?!
Training Qwen2.5-Math-7B with our Spurious Rewards improved MATH-500 by:
- Random rewards: +21%
- Incorrect rewards: +25%
- (FYI) Ground-truth rewards: + 28.8%
How could this even work⁉️ Here's why: 🧵
Blogpost: https://t.co/jBPlm7cyhr