Can multilingual text encoders borrow the semantic space from images to align their representations cross-lingually?
I am presenting my paper โCross-Lingual Representation Alignment Through Contrastive Image-Caption Tuningโ at the 4pm poster session at ACL tomorrow!
๐งต
(1/9)