Our work, "Let's Think Frame by Frame with VIP: A Video Infilling and Prediction Dataset for Evaluating Video CoT" got in EMNLP23! Congrats undergrad coauthors @vaishnavihima@DannyRose30@andy__ouyang@RyanHe88 on your 1st paper!
Stay tuned for arxiv update and dataset link!
Large language models can help understand videos by thinking frame by frame.📽️💡
A video chain of thought creates structured and unstructured scene descriptions that bridge the gap between videos and LLMs! #VideoCOT#LLM
https://t.co/yQFuKlloV9 🧵3
🚨🚨 The official edition of Foveate, Attribute, and Rationalize: Towards Physically Safe and Trustworthy AI by @sharonlevy21@WilliamWangNLP and I as part of @ucsbNLP is available now -- to appear in #ACL2023#ACL2023NLP#ACL2023Toronto! 🚨🚨
https://t.co/l8FRK0OI6h
🚀Excited to release #LLMScore, an object-centric description grounded #LLM, capable of following diverse instructions to produce the best human-correlated score with rationale for text-to-image synthesis evaluation .🧵6
📜https://t.co/iptFFBgFBN
🔗https://t.co/aGh4r20RsL
Multimodal infillings can unlock the power of computer reasoning about sequential data. 🖥️🧠
We’re excited to announce that “Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings” is available on #arxiv! #LLM#GPT#StableDiffusion#NLP#CV#VCOT 🧵1/6
Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings
Presents VCoT, a novel method that leverages chain of thought prompting with vision-language grounding to recursively bridge the logical gaps within sequential data
https://t.co/m4M5AoenEw