This is the official account of the Maieutic Lab at JHU. We broadly work on Multilingual NLP and AI. The photos are what OpenAI's DALL-E "thinks" we are...
Looking forward to visiting UVA this Friday. If you are around Charlottesville, I’d love to chat more about multilingual challenges in AI.
https://t.co/HLDvR79ECJ
The MAIEUTIC Lab (https://t.co/ejVaEe5e26) is @aclmeeting and looking for new members. If you are interested in joining, please fill out this form: https://t.co/vGkUhp6oL4 If you are in person at ACL 2026, please fill this one out as well: https://t.co/fJ8xdbF1PZ
I'll be in Seoul for #ICML2026 🇰🇷!
I'll present "DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging" at the WeightSym workshop
I'm interested in LLM efficiency, KV cache, management, continual learning,... and birding 🐦. DM if you want to chat!
🧵 The MAGMaR 2026 Shared Task results are out!
We challenged teams to retrieve relevant videos from a corpus of 110K+ multilingual clips AND generate grounded, persona-driven articles from them.
Come see the findings at our workshop at ACL on July 4!
https://t.co/9YNvHbEYKH
Excited for MAGMaR 2026 in two days. People worry about LLMs hallucinating, and often you use RAG to help mitigate that. But what if your evidence isn’t in text? We ran a whole shared task (and workshop) where you use multimodal sources instead of documents - including videos.
I'm excited to announce that this Fall I will be joining the Computer Science Department at George Mason University as an Asst. Prof. I'll be expanding my lab and looking for PhD students to work on Multilingual AI problems text, video, and speech.
https://t.co/f6PIESWRp4
Introducing "DOTResize: Reducing LLM Width via Discrete Optimal Transport-based Neuron Merging" !
We introduce an optimal transport framework for Transformer width compression that redistributes signal across neurons rather than eliminating them 🚚⚖️
🧵 1/6
Take a break from studying and test your Wordle skills with a Wheel of Fortune! tournament tomorrow from 6–8 p.m. in Room 213 of the Bloomberg Student Center! Sign up now to play against other students and the AI bots they developed over the fall semester: https://t.co/nD0vUgG3Lp
Super excited for the 1st Workshop on Multilingual Data Quality Signals (WMDQS) which is happening at #COLM2025 tomorrow. We are focused on looking all the way back to the web data that goes into all your LLMs and how we can do better at multilingual. Stop by! @COLM_conf
Congrats to Dr. Xu @fe1ixxu on a successful thesis defense of "Minimizing Language Interference for Multilingual Models". The thesis covered 9 of his first author papers encompassing:
Modeling (https://t.co/VWXCvZyUrG, https://t.co/pOIq2jpoeg, https://t.co/9HXuUhBXXz)
Introducing X-ALMA: a 50 language multilingual machine translation model. It’s average translation performance outperforms other multilingual models (including ones focused on fewer num of langs) pushing against the curse of multilinguality.
📢When LLMs solve tasks with a mid-to-low resource input/target language, their output quality is poor. We know that. But can we pin down what breaks inside the LLM? We introduce the 💥translation barrier hypothesis💥 for failed multilingual generation. https://t.co/VnrOWdNPr8
A new study from Johns Hopkins researchers @nikhilsksharma, @ZiangXiao, and @kentonmurray finds that multilingual #AI privileges dominant languages, deepening divides rather than democratizing access to information. Read more: https://t.co/Z8VAPrR3Xq
عسلامة,
As part of IWSLT 2023, we are hosting a Tunisian-to-English Speech Translation Shared Task and would love people to participate. The evaluation campaign runs April 1st-15th.
More details can be found here: https://t.co/0GbDB8cNFr
Yaishek
How to fine-tune a model when you have close to no labeled data? Check out our new paper "Language Agnostic Code-Mixing Data Augmentation by Predicting Linguistic Patterns" which introduces a zero-cost code-mixing generation method for sentiment analysis: https://t.co/FC9goOJcRI