Excited to present our work on "sharpness-aware pretraining" at ICML'26 in Seoul π°π·!
Poster #2810 | July 7, 2:00-3:45 PM | Hall A
Project: https://t.co/4V117ww5gG
Come say hi if you're interested in efficient adaptation, agent memory, or have Seoul recommendations! π
[LG] ORION: Teaching Language Models to Reason Efficiently in the Language of Thought
K Tanmay, K Aggarwal, P P Liang, S Mukherjee [Harvard University & Hippocratic AI & MIT] (2025)
https://t.co/b1biuRyS6v
@rohanpaul_ai Thanks @rohanpaul_ai for sharing our work. Our team will soon be releasing the model, data, and codebase. Excited to see the community engage with it.
The paper teaches reasoning models to think in a tiny symbolic language so they stay accurate while using far fewer tokens.
Standard reasoning models like DeepSeek R1 score well on math but write huge self talk, which makes inference slow and expensive.
To shrink this, the authors design Mentalese, where each step is a terse operator plus a tiny calculation, and they build about 40K math traces in that format.
They first fine tune small reasoning models on these traces so every solution becomes a single Mentalese script, which sharply cuts length but hurts accuracy.
Next they use reinforcement learning with a verifier that checks the final answer, sampling many candidate traces per question and rating them by reward.
Shorter Length Preference Optimization keeps correctness as the main reward but adds a small bonus when a correct trace is shorter than other correct ones, without penalizing the only correct long trace.
These choices produce the ORION models, which match strong math performance while using about 4 to 16 times fewer reasoning tokens and making training and inference several times cheaper.
----
Paper Link β arxiv. org/abs/2511.22891
Paper Title: "ORION: Teaching Language Models to Reason Efficiently in the Language of Thought"
π§΅1/10
LLMs can answer in many languages.
But do they think in them?
Even when prompted in Swahili or Thai, models often switch to English for reasoning.
This breaks interpretability and trust.
So we ask: Can LLMs reason in the input language?
New paper!π
Our work, "Ethical Reasoning over Moral Alignment: A Case and Framework for In-Context Ethical Policies in LLMs," has been accepted to the Findings of EMNLP 2023!
@monojitchou@AetherSuRa@kr_tanmay147 @0203_utkarsh
Paper Link: https://t.co/mLHQbOrZgM
Guiding Language Models of Code with Global Context using Monitors
paper page: https://t.co/0zDWrAQ9el
Language models of code (LMs) work well when the surrounding code in the vicinity of generation provides sufficient context. This is not true when it becomes necessary to use types or functionality defined in another module or library, especially those not seen during training. LMs suffer from limited awareness of such global context and end up hallucinating, e.g., using types defined in other files incorrectly. Recent work tries to overcome this issue by retrieving global information to augment the local context. However, this bloats the prompt or requires architecture modifications and additional training. Integrated development environments (IDEs) assist developers by bringing the global context at their fingertips using static analysis. We extend this assistance, enjoyed by developers, to the LMs. We propose a notion of monitors that use static analysis in the background to guide the decoding. Unlike a priori retrieval, static analysis is invoked iteratively during the entire decoding process, providing the most relevant suggestions on demand. We demonstrate the usefulness of our proposal by monitoring for type-consistent use of identifiers whenever an LM generates code for object dereference. To evaluate our approach, we curate PragmaticCode, a dataset of open-source projects with their development environments. On models of varying parameter scale, we show that monitor-guided decoding consistently improves the ability of an LM to not only generate identifiers that match the ground truth but also improves compilation rates and agreement with ground truth. We find that LMs with fewer parameters, when guided with our monitor, can outperform larger LMs. With monitor-guided decoding, SantaCoder-1.1B achieves better compilation rate and next-identifier match than the much larger text-davinci-003 model.