Introducing K2 Horizon: a connected fleet of six foundation models ranging from 0.9 billion to 375 billion parameters.
- Frontier performance: Across coding and agentic tasks, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B and 7B models setting new state of the art at their respective scales.
- Radical openness: K2 Horizon represents the largest fully open-source model launch in AI history. The fully open code, training data and recipes are a significant step forward in transparency.
Launch page: https://t.co/gg0k803SbL
Tech blog: https://t.co/g35L5xMGdS
Hugging Face: https://t.co/3Lb28JhyG9
So we're building a hard, multilingual multimodal browsing agent benchmark. Data is mostly done, just benchmarking models now. The questions are diabolical😈
Paper & data release coming soon! https://t.co/KoH6DanLHI
At ICML right now, hit me up if you are interested!
also...
🚨 ARR is recruiting extra emergency reviewers and emergency Area Chairs (ACs) for this cycle.
Emergency Reviewer Registration:
https://t.co/DtVZqCHm1o
Emergency AC Registration:
https://t.co/jKxtqzvYbX
Thank you for helping support the ARR review process.
#ARR#ACL#NLProc
Why do LLM personalization systems look good on paper… but fail in practice?
Glad to share our new preprint:
LUCid: Redefining Relevance for Lifelong Personalization
Paper: https://t.co/viAV4Mbsla
🧵
Afri-MCQA is an amazing paper if you are working in multilinguality/cultural NLP.
We found that current models still struggle a a lot when it comes to visual cultural knowledge for African regions.
Check it out!
Kudos to all the great co-authors!
Really excited about this upcoming #AAAI2026 paper led by my amazing PhD student Joan Nwatu. By reorganizing objects around what they are used for, rather than how they look, we show meaningful reductions in socioeconomic performance gaps in vision–language models.
We introduce the Culture Affordance Atlas, a culturally grounded re-annotation of Dollar Street that surfaces overlooked objects and improves model performance in lower-income contexts.
📝 https://t.co/TkiVtMNHEY
🔗 https://t.co/I32hAaDfkz
👥 Joan Nwatu @_agirlyengineer, @Longju_Bai, @OanaIgnatRo, @RadaMihalcea
PhD admissions are brutal now. You need a serious portfolio before you even start, not to mention the privilege of having connections to help you out. I just got lucky the bar was lower back then😅
This privilege is scarce, so SEACrowd is opening our Apprenticeship Program. You can work on cool AI/ML research projects guided by mentors from all over the world.
Not only will you gain hands-on experience, but this opens up connections too!
Open to people from SEA or those interested in SEA languages, especially early-career students aiming for postgrad studies!
Interested? Consider applying
I'll be in *SEM today, presenting efforts led by @AtnafuLambebo on developing technology and data for African languages. The talk also summarizes our recent survey on what we mean by low-resource in NLP. I hope to amplify the message from our group and others in our field, motivating a more nuanced discussion around low-resource languages and the ways we can make progress for understudied languages.
@emnlpmeeting
Relevant papers:
https://t.co/R2S9ENBfeS
https://t.co/wvA0OURNWj
Welcome back lunch for
RiTUAL lab: a new semester started and we have some new faces and some members completing their appointment with us. I'm thankful for the contributions and connections that the researchers in my group bring. I'm still hiring, visiting students, postdocs, short term research visits. Get in touch. See our research topics here: https://t.co/yIaTUoFGQw
Visual and auditory cues (like gaze, facial expressions, prosody) often play a significant role in understanding the interaction to answer the question correctly.
We also reduce answer-set bias using an LLM-in-the-loop process, so models must reason from context, not just guess.
Multimodal input helps, but models still struggle to integrate visual and auditory cues effectively, especially over longer contexts.
MOMENTS reveals the limitations in current systems and offers a path toward building more socially intelligent AI
Each question is grounded in social situations within self-contained stories, allowing a deeper understanding of characters and their mental states.
We go beyond beliefs/goals, covering emotions, intentions, sensory perceptions, non-literal communication, and more.
Excited to finally share MOMENTS!!
A new human-annotated benchmark to evaluate Theory of Mind in multimodal LLMs using long-form videos with real human actors.
📽️ 2.3K+ MCQA items from 168 short films
🧠 Tests 7 different ToM abilities
🔗 https://t.co/yfafwp5oIA
We’re excited to introduce CaMMT: a human-curated benchmark for culturally aware multimodal translation.
Covering 23 regions, it shows how images can help preserve cultural nuance in translation.
📷+📝=🌍
📄 https://t.co/iNxK5nSxFT
@AtnafuLambebo, @injy_hamed, @thamar_solorio
We put 5 VLMs to the test, and we found that visual input improves CSI preservation, gender marking, and lexical disambiguation, with multimodal translations often preferred by native speakers even when automatic metrics show modest gains.