Excited to see our work featured in the latest Microsoft Research Focus edition alongside other great research.
Our paper, Reasoning Beyond Labels, explores how LLMs interpret sentiment in multilingual, culturally nuanced communication and why evaluation must go beyond accuracy!
In the latest Research Focus: LLM sentiment in cultural context, robot assembly learning, smarter AI agents, verified Rust code, and CHI 2026. https://t.co/JDW828G95X
The award is here, and I had to share. 🏆
Our paper IrokoBench won the Outstanding Paper Award at #NAACL2025!
Proud to be part of this work advancing multilingual LLM evaluation across 17 African languages.
Huge thanks to @davlanade for visionary leadership.
#AfricanLanguages
🎉 Just had an incredible experience attending Conference on Neural Information Processing Systems (NeurIPS 2024)! 🎉 #NeurIPSConference - via #Whova event app
In few-shot evaluation, LLaMa 3 70B significantly benefits from few-shot examples for AfriMMLU and AfriXNLI but did not for AfriMGSM since it's only able to reason effectively on maths in English.
GPT-4o consistently improves in performance with additional few-shot examples.
GPT-4o is the best across all tasks for native African languages, however, the performance is worse for eng & fra, where GPT-4-Turbo is more than +9.0 better. Aya-101 is the best open-model, but in the translate-test setting LLaMa 3 70B is better since it's more English-centric
We provide zero/few-shot evaluation of the performance of 14 LLMs (open weight models and closed models) in two settings leveraging lm-eval:
1) in-language and
2) translate-test (where test sets were automatically translated to English using NLLB-200-3B).
We cover 18 languages in our benchmark, 16 native African languages (amh, ewe, hau, ibo, kin, lin, lug, orm, sna, sot, swa, twi, wol, xho, yor, & zul) , and two European languages (eng and fra) translated from MGSM, MMLU and XNLI datasets.