@LyxTg We also see the benefit of having more videos for a training dataset, so we plan to release our dataset creation pipeline code so future work can expand on this!
@LyxTg And since varying input frames yields different frame windows for our tasks, the total number of task examples is actually much more than 1.5k. We ensure quality of this data by carefully selecting videos and utilizing expensive data generation with human annotation.
@LyxTg Hi, thank you for your interest in our work! VIP is an inference time dataset and so we emphasize our tasks (infilling and prediction) which are frame-level tasks. Considering this, VIP has over 1.5k frames.
1️⃣To visually👀 describe each keyframe, we introduce two forms of keyframe descriptions: descriptive dense captions and FAMOuS descriptions (identifies the focus, action, mood, objects, and setting of each keyframe).
🧵3/6
3️⃣Models show weak performance on our tasks, demonstrating a need for future research in multi-hop, multi-frame video reasoning abilities of LLMS 🧐 🧵5/6
2️⃣Video Infilling Task: given n surrounding keyframes🎞️, the model is tasked to infill by predicting the p masked keyframes’ scene descriptions 🗒️
Video Prediction Task: given n preceding keyframes🎞️, the model is tasked to predict the next p keyframes’ scene descriptions🗒️ 🧵4/6
Large language models can help understand videos by thinking frame by frame.📽️💡
A video chain of thought creates structured and unstructured scene descriptions that bridge the gap between videos and LLMs! #VideoCOT#LLM
https://t.co/yQFuKlloV9 🧵3
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
Presents VCoT, a novel method that leverages chain of thought prompting with vision-language grounding to recursively bridge the logical gaps within sequential data
https://t.co/m4M5AoenEw
Multimodal infillings can unlock the power of computer reasoning about sequential data. 🖥️🧠
We’re excited to announce that “Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings” is available on #arxiv! #LLM#GPT#StableDiffusion#NLP#CV#VCOT 🧵1/6