🌍 LLMs can use long chain-of-thought (CoT) to reason in English, but what about other languages?
New paper w/ @BerkeleyNLP: We study how scaling, pretraining, post-training & inference affect long CoT across 9 languages.
Spoiler: English long CoT ≠ multilingual long CoT 🧵
Please join us in congratulating Kyle Mahowald (@kmahowald) on his promotion to Associate Professor, effective August 16th, 2026. Through his research and teaching, Kyle has made enormous contributions to our department, the university, and the field. Congratulations, Kyle!
We introduce SP3F: an algorithm to improve LLM reasoning in lower-resource languages (e.g., Indonesian, Bengali, and Swahili), no manual data collection needed! SP3F produces models that out-perform fully post-trained models across multiple tasks and languages🧵 [1/N]
🌍 LLMs can use long chain-of-thought (CoT) to reason in English, but what about other languages?
New paper w/ @BerkeleyNLP: We study how scaling, pretraining, post-training & inference affect long CoT across 9 languages.
Spoiler: English long CoT ≠ multilingual long CoT 🧵
Delighted Sasha's work using mech interp to study complex syntax constructions won an Outstanding Paper Award at EMNLP!
And delighted the ACL community continues to recognize unabashedly linguistic topics like filler-gaps, and the huge potential for LMs to inform such topics!
Finally, here are some other great works in this space covering inference-time scaling, language mixing, and language compliance.
Yong et al: https://t.co/vmXcyFdTbI
Son et al: https://t.co/sgEz7VccNf
Wang et al: https://t.co/IThpJMYpvd
Qi et al: https://t.co/A6FTB253Sj
🌍 LLMs can use long chain-of-thought (CoT) to reason in English, but what about other languages?
New paper w/ @BerkeleyNLP: We study how scaling, pretraining, post-training & inference affect long CoT across 9 languages.
Spoiler: English long CoT ≠ multilingual long CoT 🧵