This idea inherently applicable to Image/Video and other modalities while more efficient, flexible, and scalable compared to ViT. thanks to Sujoy Paul, @gaganjain1582 , @vtresp , and @jainprateek_
Excited to announce that our
@GoogleDeepMind
paper, "LookupViT: Compressing visual information to a limited number of tokens," has been accepted at
@eccvconf
three key takeaways...
https://t.co/Vu2V2GMVN7
- Introducing high-resolution lookup tokens and smaller compressed tokens.
- Expensive self-attention and MLP operations are restricted on the compressed tokens.
- Updated information between compress and lookup tokens interact through a shared bidirectional cross-attention.
If you are in Washington D.C. or attending @RealAAAI , check out our poster on 12th December named InstanceFormer, it is a simplified, tracker-free, single-stage video instance segmentation framework for long and occluded videos.@hannan_tanveer@sahandsharif
#ECCV2022 presenting now โRelationfomerโ it can generate insightful object-relation graph out of any image ..visit our poster in 36th counter..see u there ,paper: https://t.co/TSTJ04ppIU
Delighted to be at @eccv2022 in Tel Aviv; if you are interested in a scene or object-relation graph from natural to medical image, please visit the poster "Relationformer" on the 27th at Hall B in the first half.@supro_s@sahandsharif@vtresp Full paper: https://t.co/TSTJ047Okk
Delighted to share our paper 'Relationformer: A Unified Framework for Image-to-Graph Generation.' got accepted in #ECCV2022. It is a single-stage "object- relation" generation method for scene graphs, 3D vessels, and road networks. Thanks to @supro_s ,@sahandsharif ,@vtresp