Update for our latest work: a Text2Motion model based on MMDiT and Flow Matching, outperforming current SOTA like NVIDIA’s GENMO/Kimodo and Tencent Hunyuan’s HY-Motion. Welcome to check it out and try our demo!
Project: https://t.co/YGew4qEAfT
Demo: https://t.co/aCLWB8N2ze
#ZJULifeBlog: Have you tried trail walking from Yuquan Campus to Zhijiang Campus? Check out the No.5 activity on ZJUers’ must-do List at ZJU: https://t.co/NszclwOCw5
#Campuslife#StudyatZJU
CustomVideo: Customizing Text-to-Video Generation with Multiple Subjects
paper page: https://t.co/joFsWig1Bf
Customized text-to-video generation aims to generate high-quality videos guided by text prompts and subject references. Current approaches designed for single subjects suffer from tackling multiple subjects, which is a more challenging and practical scenario. In this work, we aim to promote multi-subject guided text-to-video customization. We propose CustomVideo, a novel framework that can generate identity-preserving videos with the guidance of multiple subjects. To be specific, firstly, we encourage the co-occurrence of multiple subjects via composing them in a single image. Further, upon a basic text-to-video diffusion model, we design a simple yet effective attention control strategy to disentangle different subjects in the latent space of diffusion model. Moreover, to help the model focus on the specific object area, we segment the object from given reference images and provide a corresponding object mask for attention learning. Also, we collect a multi-subject text-to-video generation dataset as a comprehensive benchmark, with 69 individual subjects and 57 meaningful pairs. Extensive qualitative, quantitative, and user study results demonstrate the superiority of our method, compared with the previous state-of-the-art approaches.
Mobile ALOHA's hardware is very capable. We brought it home yesterday and tried more tasks! It can:
- do laundry👔👖
- self-charge⚡️
- use a vacuum
- water plants🌳
- load and unload a dishwasher
- use a coffee machine☕️
- obtain drinks from the fridge and open a beer🍺
- open doors🚪
- play with pets🐱
- throw away trash
- turn on/off a lamp💡
Project website: https://t.co/9rzIX8wLEp
Co-lead @tonyzzhao, advised by @chelseabfinn
(amazing photographing from @qingqing_zhao_ )
Some predictions for 2024 – keeping only the more controversial ones. You certainly saw the non-controversial ones (multimodality, etc) already
1. At least 10 new unicorn companies building SOTA open foundation models in 2024
Stars are so aligned:
- a smart, small and dedicated team can reach close to OpenAI level in a few months as Mistral, https://t.co/ElVO5V4Bg6 or https://t.co/aqWGInU1AP showed
- non-AI startups are struggling to raise money while VCs are eager to join the AI revolution
- knowledge around training large models keeps spreading from teams to teams each time a new SOTA model is created leading many new venture to form
- consequence: we now understand much better the push for regulatory capture that we witnessed from early AI-labs in 2023 (e.g. the idea of licences to train models, etc)
2. 2024 will be reality-check for older AI-unicorn pioneers
Optimistic bet: we'll see a few older AI unicorn startups successfully transitioning to sustainable business models by leveraging widespread public adoption of AI
3. Model quality will be harder and harder to evaluate in 2024
With the surge of models and the saturation of open-benchmarks, users will tend to fall back on "brand quality perception". It will become essential for new teams to not be perceived as cheating on public evals and leaderboards
4. The return of academia
Academia is back as we saw at NeurIPS 2023. With many private and open-source labs closing the doors on publishing their results and data, academia rise again in visibility and is shining with many impactful papers in 2023 and exciting new work coming
5. Dangerous times for annotation companies
It's much easier to quickly spin and iterate on a pay-by-usage API than to hire and manage annotators. With model performance strongly improving and the privacy guarantee of open models, it will be harder and harder to justify making complex annotations contracts.
6. The rise of synthetic data
We're running out of human data, there are many copyright questions on these and our largest models are already reaching human level annotation on many tasks. The next step coming is about large and good quality synthetic data.
I spent over 500 hours in the last 6 months mastering ChatGPT.
Using prompt frameworks is by far the most effective way to level up your outputs.
But, most don't know where to start.
So, I made a Cheat Sheet to help you maximize ChatGPT using these 5 prompt frameworks:
ZJU's international campus has committed to building itself into a world’s leading sustainable campus. With efforts of past few years, the campus received the ISO14001:2015 certification for its sustainable management capacity.
https://t.co/F9PQ0W8fmk
#ZJUI#Z4G
Can you believe the photos below are all inspired by buildings on ZJU campus? Students attending the course "Architecture Photography" shared their wild imagination and unique perspectives in their assignments: https://t.co/0mJXocd0cy
Photo: 浙大官微
#ZJUer#architecturelover
What comes to your mind when you think of #Chinesepaintings? Instead of painting landscapes, MA Nan, associate professor at ZJU's School of Art and Archaeology sketches portraits of her students to teach them basic Chinese painting techniques.
Photo: 浙大官微
#ZJUer
#HepatitisC virus drug simeprevir could be a potent treatment for #COVID19 which can target two major viral proteins, according to our global study with @hkumed in @ACSCentSci: https://t.co/L4tnp7Ul1L
Tweetorial by Prof Billy NG (@wai_lung_ng): https://t.co/xp2VbSYRAi