Thanks @_akhaliq for the sharing. Check out our MVHumanNet, the largest-to-date dataset of multi-view human captures, with 4,500 human identities and 9,000 daily dressing. We plan to release it in the next months. Following our MVImgNet, hope it can help.
MVHumanNet: A Large-scale Dataset of Multi-view Daily Dressing Human Captures
paper page: https://t.co/kOFmiKl2TP
In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while remarkable progress has been made with models trained on large-scale synthetic and real-captured object data like Objaverse and MVImgNet, a similar level of progress has not been observed in the domain of human-centric tasks partially due to the lack of a large-scale human dataset. Existing datasets of high-fidelity 3D human capture continue to be mid-sized due to the significant challenges in acquiring large-scale high-quality 3D human data. To bridge this gap, we present MVHumanNet, a dataset that comprises multi-view human action sequences of 4,500 human identities. The primary focus of our work is on collecting human data that features a large number of diverse identities and everyday clothing using a multi-view human capture system, which facilitates easily scalable data collection. Our dataset contains 9,000 daily outfits, 60,000 motion sequences and 645 million frames with extensive annotations, including human masks, camera parameters, 2D and 3D keypoints, SMPL/SMPLX parameters, and corresponding textual descriptions. To explore the potential of MVHumanNet in various 2D and 3D visual tasks, we conducted pilot studies on view-consistent action recognition, human NeRF reconstruction, text-driven view-unconstrained human image generation, as well as 2D view-unconstrained human image and 3D avatar generation. Extensive experiments demonstrate the performance improvements and effective applications enabled by the scale provided by MVHumanNet. As the current largest-scale 3D human dataset, we hope that the release of MVHumanNet data with annotations will foster further innovations in the domain of 3D human-centric tasks at scale.
LRM: Large Reconstruction Model for Single Image to 3D
paper page: https://t.co/5lvfMVZT3D
propose the first Large Reconstruction Model (LRM) that predicts the 3D model of an object from a single input image within just 5 seconds. In contrast to many previous methods that are trained on small-scale datasets such as ShapeNet in a category-specific fashion, LRM adopts a highly scalable transformer-based architecture with 500 million learnable parameters to directly predict a neural radiance field (NeRF) from the input image. We train our model in an end-to-end manner on massive multi-view data containing around 1 million objects, including both synthetic renderings from Objaverse and real captures from MVImgNet. This combination of a high-capacity model and large-scale training data empowers our model to be highly generalizable and produce high-quality 3D reconstructions from various testing inputs including real-world in-the-wild captures and images from generative models.
Given the performance of SAM on 2D images, do we still need training 3D understanding models? Check out our SAMPro3D for segmenting any 3D indoor scenes: https://t.co/OlA9nwuZ7U. It gets impressive results which are often better than human annotations.
Happy to announce: A series of our sketch-based character modeling systems have been open-sourced: Animal Head Modeling (UIST 21 https://t.co/s0nJ5MSJ3B), Generic Character Head (TVCG 23 https://t.co/eLSx0CNXzs), and Biped Cartoon Characters (CVPR 23 https://t.co/FmBdI3eQUy).
We maintained a webpage, "CV-Highlight-Papers", which contains all Oral ("Highlight" at CVPR 2023) papers in CVPR/ICCV/ECCV/NeurIPS/ICLR from 2017 to now. And each paper owns github link, project page link, stars and citations. Hope it can help.
https://t.co/TCDgZJHDET
Excited to share our new work "Learning 3D Scene Priors with 2D Supervision" #CVPR2023!
3D labels are costly! We learn priors of object semantics and shapes in 3D scenes with only 2D supervision.
https://t.co/wrCqzqiGJV
https://t.co/dkssxMV8sy
@yinyu_nie@angelaqdai@XiaogHan
RaBit: Parametric Modeling of 3D Biped Cartoon Characters with a Topological-consistent Dataset
abs: https://t.co/sn5nfbZsDf
project page: https://t.co/GCRTpPJOK0
RaBit: Parametric Modeling of 3D Biped Cartoon Characters with a Topological-consistent Dataset
abs: https://t.co/sn5nfbZsDf
project page: https://t.co/GCRTpPJOK0
Happy to share our work on 3D Hair modeling, accepted by #CVPR2023
arxiv: https://t.co/oXzMneKGtu
project page: https://t.co/d5faSU46fm
HairStep -- Supplementary Video https://t.co/uZRQ3JIZkk via @YouTube
Happy to share our work at #CVPR2023. Now, you can create cartoon characters easily.
arxiv: https://t.co/PQ4lVJo3sN
Project page: https://t.co/UcAGB9RkYa
Excited to share again our work MVImgNet at #CVPR2023.
arxiv: https://t.co/g6I483W5xJ
Project page: https://t.co/SKBleWMqDn
MVImgNet, CVPR 2023 https://t.co/7CUHQjxrjL via @YouTube
Happy to share our work on 3D Hair modeling, accepted by #CVPR2023
arxiv: https://t.co/oXzMneKGtu
project page: https://t.co/d5faSU46fm
HairStep -- Supplementary Video https://t.co/uZRQ3JIZkk via @YouTube
Thanks to
@_akhaliq
for sharing our recent work, "MVImgNet: A Large-scale Dataset of Multi-view Images", accepted by #CVPR2023. Data will be released soon.😃