Our Geo Regional Asia group is hosting a panel September 4th with @roopalgarg , @bka3 and @AndreaDBurns ! Be sure to add it to your calendar! 🧠
Learn more: https://t.co/MLnCxrUwKz
A belated announcement that my final work from my PhD was accepted to ACL Findings! I’ll be presenting on Tell Me What’s Next: Textual Foresight for Generic UI Representations
A new approach to learning UI features that has fully open sourced code, pretrained models, and data!!
To train our model we build a new dataset on top of prior work, OpenApp, which enables several other baselines. Code, models, and data are made available!! Check it out
Congrats to the team for advancing hyper detailed image descriptions!! Grateful to have been a part of this project - check out the released data on GitHub or Huggingface 🥳🥳
#NeurIPS2023 bound tomorrow. Looking forward to discovering new papers and having fun conversations. DM if you'd like to chat!
On Thu (5-7 PM CST), I'll help present our poster
"Learning Human Action Recognition Representations Without Real Humans" https://t.co/8br4A9By42
Come chat with me about it on the 8th at 4pm. Also happy to meet for coffee or lunch and make more NLP friends!! say hi 👋📷 and we can talk about vision-language tasks, representation learning, or digital domains like webpages and mobile apps
I'm presenting our WikiWeb2M work on webpage understanding at @emnlpmeeting this Friday - have you ever wanted text, image, and structure relating them all in one place? we provide it for the first time in a unified multimodal webpage sample 🧵
Prefix Global can prioritize a subset of webpage inputs to specially attend to the rest, while less important image or text has local attention. This is only made possible by having our webpage structure - no more bag of images or words!!