[2/2] GPT-6 Astra brings memory and reasoning to mobile manipulation. 🚀🚀🚀
It learns new behaviors in context from human videos, spots fine details, understands complex scenes, interacts naturally and remembers past interactions.
Embodied intelligence in motion. 👇
[1/2] GPT-6 Astra might be the craziest robot policy I’ve seen. 🤯
Give it ONE demo per task, and it can control a real robot to unscrew a bottle cap, insert a plug into a power strip, clear obstacles, and even imitate a human dance.
One context. New behavior. 👇
📢Introducing Generated Reality📢
A world model for XR that turns your tracked hand and head poses into an interactive, generative video experience. Take world models to the next level by interacting with the world using your own body!
🔗https://t.co/bgDDO8Laix
1/4
🏆 The 3D Generation Leaderboard is LIVE!
We introduce a novel standardized benchmark for 3D generation models, dubbed Hi3DEval.
💥 30 existing methods are already on board!
📄 Paper: https://t.co/YWfFO3JwGT
🌐 Explore on Hugging Face:
https://t.co/7QB5MlgUV2
The context size of video world models is only a few frames. Like a human with severe memory loss! We design a long-term memory for world models based on explicit 3D representations inspired by the human mind. This enables long-term consistency. https://t.co/TExH0r78xV
1/3
💡RelightVid: Temporal-Consistent Diffusion Model for Video Relighting.
A temporally consistent video relighting framework. By extending IC-Light with temporal layers and multi-modal conditions, it enables coherent and flexible video relighting.🚀https://t.co/gnHu4Tc1Db
🎉 Excited to introduce IDArb! 🎉
Our method can predict plausible and 𝗰𝗼𝗻𝘀𝗶𝘀𝘁𝗲𝗻𝘁 geometry and PBR material for 𝗮𝗻𝘆 𝗻𝘂𝗺𝗯𝗲𝗿📷 of input images under 𝘃𝗮𝗿𝘆𝗶𝗻𝗴 𝗶𝗹𝗹𝘂𝗺𝗶𝗻𝗮𝘁𝗶𝗼𝗻𝘀☀️ !
Webpage: https://t.co/GvfyvbEq25
🚀 We’re excited to announce the release of InternLM-XComposer2.5-OmniLive (IXC2.5-OL), a comprehensive multimodal system designed for long-term streaming video and audio interactions. This fully open-sourced project delivers functionality similar to Gemini 2.0 Live Streaming and OpenAI Her, with standout features including:
🎥 Chat with Streaming Video & Audio
💾 Long-Term Memory for recalling past video experiences
🏆 Competitive Performance across various video and audio perception benchmarks
📄 Paper: https://t.co/oiBfU74lR4
💻 Code: https://t.co/hP881kVnlM
📦 Models: https://t.co/t7hGEuj9Qx
✨ Immerse yourself in multimodal interaction and create your own app today!
#Gemini2 #OpenAI #ChatGPTAdvancedVoice
😻Fine-Grained Visual Attributes for GenAI😻
#NeurIPS2024 🍎FiVA🍊 is a fine-grained visual attributes dataset and a framework that decouples different visual attributes for GenAI
- Project: https://t.co/hhSlc7PFQm
- Code: https://t.co/Ggji0AluDN
- Data: https://t.co/LgRjvcShl1
🚀Excited to introduce 𝗜𝗺𝗮𝗴𝗶𝗻𝗲𝟯𝟲𝟬, the 1st framework that lifts standard videos into 360 videos with rich and structured motion, unlocking dynamic scene experience from full 360 degrees. 🎥🌍
- Webpage (w/ VR): https://t.co/0YsLSEAAT4
- Arxiv: https://t.co/0nBUyB9AKs
Excited to see this incredible progress in 3D panoramic scene generation! 🚀 We also have a preliminary attempt in this direction. Take a look and share your thoughts! 🙌
- LayerPano3D Paper: https://t.co/0er9iV1Na1
- Project page: https://t.co/p8PmzVZMyf
We are excited to introduce LayerPano3D, a novel framework to generate full-view, explorable panoramic 3D scene from a single text prompt! More examples are in our project page.
✨Project: https://t.co/p8PmzVZeIH
✨Paper: https://t.co/0er9iV1fkt
✨Code: https://t.co/LEFIiUE69i
🤩3D Generation Arena🤩
Lots of 3D generation models come out recently. But which one is preferable from human perception?
** Welcome to play with ~20 #3DGen models in our arena with both text-to-3D and image-to-3D @_akhaliq
- 3DGen-Arena @huggingface : https://t.co/bF2sTzWXxn
📢📢Excited to release 3DGen-Arena, an open 3D Benchmarking platform.
⚔️Two tracks: Text-to-3D & Image-to-3D.
🎯Nineteen models: 9 for Text & 13 for Image.
🏆The Leaderbord is waiting for your votes!
Let's play with 3D models and vote at https://t.co/R8ivy7zJVR!
AI will make creating materials a breeze for 3D artists!
Make-it-Real utilizes GPT-4V to recognize and describe materials, allowing the construction of a detailed material library.
https://t.co/SR8nDcSzLH
ComboVerse
Compositional 3D Assets Creation Using Spatially-Aware Diffusion Guidance
Generating high-quality 3D assets from a given image is highly desirable in various applications such as AR/VR. Recent advances in single-image 3D generation explore feed-forward models
Wanna evaluate your own "dreamer" but facing the challenge of expensive user studies? Dive into our newest work for #CVPR2024: leveraging GPT-4V for a concise, versatile, and human-aligned evaluation. Give it a shot!
Looking for a way to evaluate your text-to-3D model?
We found that GPT-4V can be prompted to be a human-aligned and versatile 3D evaluator!
Arxiv: https://t.co/f3W9YkXfUX
Code: https://t.co/RC7oAVlsyj
Page: https://t.co/JhBzpkYi7S
(1/2) We are actively seeking PhD candidates from various countries to foster diversity in our research group at Nanyang Technological University. Know someone interested in a PhD with us? Please refer them to our team. Thanks for supporting diversity in academia! 🌍🎓
Gemini vs GPT-4V: A Preliminary Comparison and Combination of Vision-Language Models Through Qualitative Cases
paper page: https://t.co/GwqhvEMJlW
The rapidly evolving sector of Multi-modal Large Language Models (MLLMs) is at the forefront of integrating linguistic and visual processing in artificial intelligence. This paper presents an in-depth comparative study of two pioneering models: Google's Gemini and OpenAI's GPT-4V(ision). Our study involves a multi-faceted evaluation of both models across key dimensions such as Vision-Language Capability, Interaction with Humans, Temporal Understanding, and assessments in both Intelligence and Emotional Quotients. The core of our analysis delves into the distinct visual comprehension abilities of each model. We conducted a series of structured experiments to evaluate their performance in various industrial application scenarios, offering a comprehensive perspective on their practical utility. We not only involve direct performance comparisons but also include adjustments in prompts and scenarios to ensure a balanced and fair analysis. Our findings illuminate the unique strengths and niches of both models. GPT-4V distinguishes itself with its precision and succinctness in responses, while Gemini excels in providing detailed, expansive answers accompanied by relevant imagery and links. These understandings not only shed light on the comparative merits of Gemini and GPT-4V but also underscore the evolving landscape of multimodal foundation models, paving the way for future advancements in this area. After the comparison, we attempted to achieve better results by combining the two models. Finally, We would like to express our profound gratitude to the teams behind GPT-4V and Gemini for their pioneering contributions to the field. Our acknowledgments are also extended to the comprehensive qualitative analysis presented in 'Dawn' by Yang et al. This work, with its extensive collection of image samples, prompts, and GPT-4V-related results, provided a foundational basis for our analysis.
I’m recruiting multiple PhD students in my group at @CIS_Penn in the following areas:
- Neural Representations and rendering for 3D/4D Reconstruction
- 3D Generative Models
- Human Motion Generation
- LLM guided Graphics and Vision
- Neural Representations for Robotics
etc.