π Absolutely thrilled to share that our paper, "Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis" received the π Resource Award π at #NAACL2024!
you can say the authors of <BLEND: A Benchmark for LLMs on Everyday Knowledge in Diverse Cultures and Languages> were dedicated by looking at Table 1.
This paper got the best non-archival paper award! π
Congratulations to @JunhoMyung00211, @nlee0212 and @jodieyzhou who brilliantly led the project, and to all the many collaborators involved in the project! ππΌππΌ
π Absolutely thrilled to share that our paper, "Exploring Cross-Cultural Differences in English Hate Speech Annotations: From Dataset Construction to Analysis" received the π Resource Award π at #NAACL2024!
I would like to thank my co-authors, Chani Jung,
@JunhoMyung00211, @jin__jiho, @CamachoCollados, @imjuhokim, and my advisor, @aliceoh, for their unlimited support in making this happen. π
π This research highlights the need to understand and embrace cross-cultural variances to build fairer and more inclusive AI. By enhancing how we evaluate and adapt hate speech models and datasets, we aim to promote more respectful and inclusive online interactions and beyond.
@CamachoCollados@aliceoh@imjuhokim@jin__jiho Huge thanks for all your advice on this research!! It was such a great opportunity working with you. (+ our BLEnD project too π)
Really fun paper on multilingual multicultural everyday commonsense benchmark as a result of an incredible collaboration with @CamachoCollados@nedjmaou, led by @jodieyzhou @JunhoMyung00211 and @nlee0212.
What do people in North Korea eat while watching sports? What do kids do after school in Ethiopia? The appendices of the paper contain really interesting examples, and the GitHub repo contains all Qs and human annotations.
LLMs are of course quite good at these everyday commonsense Qs for US and a couple other highly represented cultures, and much worse in low-resource languages and unfamiliar cultures. (multilingual) LLMs actually answer these Qs for under-represented cultures better in English than the local languages!
This is an updated version of the same paper to be presented at @c3_nlp in August as part of #acl2024. Hope to see you there!
https://t.co/HSPwwtXzET
π€ We also evaluated current LLMs on our dataset to check for cultural biases in English hate speech classification. While these models show high label agreement from labels from Western countries, they struggle to make culturally tailored predictions when prompted to do so.
We introduced the CREHate dataset, a robust collection of 1,580 samples from 5 countries, which highlighted that while there's strong agreement in labels among the Anglosphere, the variations grow with cultural distance.
Itβs truly grateful to see our work on the cross-cultural aspects and the inclusiveness of hate speech recognition earn the 'Resource Award', as this acknowledgment reaffirms the community's recognition of these issues as crucial areas for research!
My student @nlee0212 will be at #naacl2024 to present our paper on "Exploring Cross-Cultural Differences in English Hate Speech Annotations" https://t.co/9rYKTCWmQl
We start with the simple fact that English is widely spoken globally, but current NLP datasets are focused on how English is used in North America.
This raises a particularly important concern for tasks like hate speech detection, where many cultural and regional nuances matter. We collected social media posts containing hate speech from πΊπΈπ¬π§πΈπ¬π¦πΊπΏπ¦, and we asked annotators in those countries to annotate the union of all regional posts.
There is low agreement among the annotators from the different countries, and LLMs perform worse for non-US annotations, especially Singapore and S. Africa.
This highlights the importance of considering the cross-cultural differences in NLP research, especially when building benchmarks and evaluating LLMs. With @imjuhokim@CamachoCollados@jin__jiho