Thrilled that our new multilingual reasoning benchmark, Macaron, was accepted to ACL! It's a template-based, human-written benchmark spanning 20 languages.
Attending ACL? Stop by our poster at 11:00 in Poster Session G.
Paper: https://t.co/5IReJte4Np
Unfortunately, I'm not attending ACL 2026, but if you are into competitive programming and coding benchmarks in general, please stop by our poster and say hi to my students today (July 7th).
Session 16 - Poster Session F, 9:00 AM
Paper: https://t.co/i7Kg9aXQM3
Happy to share that We will be presenting this at ACL 2026 in San Diego!
Findings of ACL · Poster Session F · Tuesday, July 7 · 9:00–10:30 AM · Grand Hall
Come chat if you're there!
(1/9) Excited to share our new paper🥳,
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming.
very excited to announce token order prediction has been accepted into ICML 2026!!
see you in Seoul!
very happy with how the paper turned out. we have an updated version of the paper on arxiv, and we’ll have a camera-ready version soon as well: https://t.co/cMOqnRjIFa
Happy to have this work accepted in ICML 2026 🥳
An arguably simple extra loss that learns the order of upcoming tokens as an alternative of MTP.
See our updated paper for more results and discussion: https://t.co/zid62pmOww
See you in ICML
Should we treat LLM benchmarking like an annual Olympiad event?🏆
With current benchmarks, it is too easy to overfit tasks or manipulate settings. In some cases, people just cheat / being narrow-tuned to a specific benchmark (*cough* LLaMa-4)
What if we organized an annual, Olympiad-like event? The tasks must be sealed and unknown. Models cannot study for the test. They must be prepared for anything. We explain this in our new position paper.
I am an IOI alum long time ago. I practiced for years to master many algorithms. I wanted to be ready for whatever appeared on the contest day. I believe general LLMs should face the same standard. If they are truly general, they should be ready for whatever use cases.
We propose a flow similar to how we typically organize an Olympiad:
- Call for Task: We propose an open solicitation for challenging, high-quality tasks from the global research community.
- Organizing Committee: A dedicated team curates and improves these submissions. They verify task quality and diversity.
- Model Developers: Developers submit their systems blindly before the tasks are revealed. This prevents teams from iterative gaming or manual tuning once the exam starts.
- The Actual Olympiad: Evaluation happens in a synchronized, short window. The sealed tasks are released, and all models are tested simultaneously to maintain total integrity under the same setting.
Once it is done, everything will be released for reproducibility.
Read the full position paper here: https://t.co/FKCWW9x7V6
We worked on this together with my student @jcblaisecruz
Let me know your thoughts!
New position paper!
📄 "LLM Olympiad: Why Model Evaluation Needs a Sealed Exam"
We argue that NLP needs an Olympiad-style event: seal the problems, freeze submissions, run one harness, release everything for audit.
w/ @AlhamFikri
Paper: https://t.co/0uq3LlvImd
In competitive programming, the post-contest discussion is mainly about the algorithm, not the code
It’s (mainly) a problem-solving contest, yet LLMs are often benchmarked only as coders
We revisited this by evaluating the editorial and reasoning vs end-to-end code generation
@SamaHadhod's paper exposes how LLMs struggle with code even after nailing the plan in competitive programming tasks. New editorial-centric benchmark + released ICPC-style problems. Very very cool ngl
(9/9) Many thanks to my co-authors: Alaa Elsetohy, @fredyhudi, @jcblaisecruz, @Steven7Halim, and @AlhamFikri.
Happy to chat if you’re thinking about reasoning evals for code or CP-style tasks.
(1/9) Excited to share our new paper🥳,
Idea First, Code Later: Disentangling Problem Solving from Code Generation in Evaluating LLMs for Competitive Programming.
(8/9)
We release 83 ICPC-style problems with expert-written gold editorials and full official test suites.
Project page: https://t.co/ZCsBOYFvSa
arXiv: https://t.co/Y10JfucMyR
Dataset: https://t.co/yAy1DUdomJ
1/11 Proud to share our new paper: SENSIA (SENse-based Symmetric Interlingual Alignment) — a sense-based approach to multilingual adaptation.
Goal: explicit representation-level alignment of meaning.
Most culture test benchmark is mostly static, which may lead to data saturation and leakage, hence making the score not reliable to measure the capability of LLMs.
Thus, we benchmark these LLMs to play a social deduction game!