Introducing Ego2Web from Google DeepMind and UNC Chapel Hill, accepted to #CVPR2026.
AI agents can browse the web. But can they act based on what you see? Existing benchmarks focus only on web interaction while ignoring the real world.
Ego2Web bridges egocentric video perception and web execution, enabling agents that can see through first-person video, understand real-world context, and take actions on the web grounded in the egocentric video.
This opens a path toward AI assistants that operate seamlessly across physical and digital environments. We hope Ego2Web serves as an important step for building more capable, perception-driven agents.
🧵👇
Introducing Long-LRM++ — for feed-forward, high-res, detail-preserving scene reconstruction
✨ Up to 64 960×540 inputs
🔍 Readable text
📉 4× fewer Gaussians
⚡ Real-time rendering
📷 End-to-end from unposed inputs w/ DA3 poses in 11s (w/⬆️ quality than DA3’s own GS predictor ;)
Do we really need massive curated 3D scene data for interactive world generation?
#SAM3D, #WorldGen say yes.
We say no.
I-Scene learns better spatial knowlesge using only 25K randomly composed instances.
🔑 Key insight:
We reprogram the instance generator to infer support, proximity, and symmetry from purely geometric cues for generating interactive scenes.
🧠 Scene-context attention
👁️ View-centric space
🧱 Random composition beats expensive curation
🌐 https://t.co/AMNThlv0NT
💻 https://t.co/nA8HICjDwz
🧵 Details below [1/6]
CSRankings counts publication in top conferences to rank professors/universities. But this encourages researchers to pursue quantity rather than quality.
We propose https://t.co/uDZLqYkD1g, a new university ranking system that tries to measure quality instead of quantity of publications.
How can we measure the quality of the publications? We believe that 1) The quality of research is best understood and evaluated by peers in the same research area;
2) With careful and informed use, LLMs can reveal the implicit quality judgments that peers convey through their citation practices and writing across large volumes of scholarly work.
Hence, we developed the new ranking system where we analyze research papers from major AI conferences with LLMs.
For each paper, we ask an LLM what are the 5 most important papers to this paper. In other words, the five works that most strongly influence the study. By doing this, we trace which papers and authors are consistently seen as inspirational and foundational to new discoveries in the field.
We ran the model on all papers from top conferences in machine learning, computer vision, natural language processing and information retrieval from 2020 - 2025, and filtered references to only have those from 2000 onwards.
Next, we map these influential authors to their affiliated universities using the CSRankings name–affiliation database. Each time a paper is recognized as one of the “top five references” in another work, its authors and their institutions receive credit. To keep the scoring fair, points are divided by the number of co-authors, ensuring balanced recognition across collaborations.
The result is a new kind of academic ranking: one that rewards universities not just for publishing often, but for producing research that endures, inspires, and drives the field forward. This approach highlights scholarly influence and provides students, researchers, and institutions with a clearer picture of where the most impactful work is happening.
Note that we believe that CSRankings had substantially improved university rankings in computer science by replacing subjective, reputation-based measures, such as those in US News, with more objective indicators, but the LLM era allows us to do something potentially better!
Due to computational resource limits, we were only able to run it with a small 7B language model. It is also a project primarily led by undergraduate and master students from Oregon State University and University of California Santa Cruz. As a result, the system is very much a work in progress and will inevitably contain errors and blind spots. We actively welcome community feedback, new collaborators and contributions of GPU compute so that we can run larger LLMs, obtain more reliable results and improve the methodology.
Long-LRM will be presented tomorrow at #ICCV2025 Poster Session 1 (11:30 AM) as a Highlight Paper!
🚀The first generalizable GS–based approach for high-res, wide-coverage 3D reconstruction in 1 second.
Come check it out & chat with us!
🧩Code & weights: https://t.co/ZsFqoh8gq2
Can we scale 4D pretraining to learn general space-time representations that reconstruct an object from a few views at any time to any view at any other time?
Introducing 4D-LRM: a Large Space-Time Reconstruction Model that ...
🔹 Predicts 4D Gaussian primitives directly from multi-view tokens (no motion vectors, no HexPlane);
🔹 Uses a clean, minimal Transformer backbone;
🔹 Generalizes with fast, high-quality feedforward rendering at any view and infinite frame rate.
Check out more interactive demos and scaling behaviors on our homepage/paper.
👉Website: https://t.co/DSjVVpI37n
👉Paper: https://t.co/VmMUrffcyp
💥 Think more real data is needed for scene reconstruction? Think again!
Meet MegaSynth: scaling up feed-forward 3D scene reconstruction with synthesized scenes. In 3 days, it generates 700K scenes for training—70x larger than real data!
✨ The secret? Reconstruction is mostly non-semantic! No need to rely heavily on real or highly realistic synthetic data.
🌐 Project: https://t.co/zsOQYwvpiK
(1/4)
Novel view synthesis has long been a core challenge in 3D vision. But how much 3D inductive bias is truly needed? —Surprisingly, very little!
Introducing "LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias"—a fully transformer-based approach that enables scalable, generalizable, and fully data-driven novel view synthesis, from sparse posed inputs. 🧵(1/6)
Project Page: https://t.co/Alqo3s0Qjt
Hate waiting 10 minutes for 3D GS to render your favorite indoor or outdoor scenes? ⏳ Our feed-forward solution, Long-LRM, cuts it down to just 1 second! ⚡️ With a straightforward mix of Mamba2 and transformer, it scales up to 32 high-res input images. https://t.co/brAgawmtV3