Can multilingual text encoders borrow the semantic space from images to align their representations cross-lingually?
I am presenting my paper “Cross-Lingual Representation Alignment Through Contrastive Image-Caption Tuning” at the 4pm poster session at ACL tomorrow!
🧵
(1/9)