LLMs that reason via Chain-of-Thought often keep thinking long after they've already found the answer, wasting tokens and compute. We built Terminator🤖to fix this: a lightweight probe that predicts when an LRM has generated its final answer and exits the CoT early.
Despite unanimous acceptance among the reviewers, our work was sadly rejected from the main conference. Since then, we have made several improvements to our arXiv paper, which are also available in the workshop version.
https://t.co/SrfFRtIig4
I'm excited to be attending ICML this year. Please come check out our work, TERMINATOR, a method for early exiting during LLM reasoning, at the AdaptFM workshop on July 11! 🇰🇷
Happy to chat!
I’ve been working through an exciting idea, but first I needed to unpack the machinery behind GRPO, starting from REINFORCE.
I’ve made my notes public; the link is in the reply below.
After giving my advisor a presentation on an earlier version of these notes, she thought it was great and encouraged me to record a YouTube video to share it.
And you know what, maybe I will do exactly that!! But first I need to attend to my research idea…
I’ve been working through an exciting idea, but first I needed to unpack the machinery behind GRPO, starting from REINFORCE.
I’ve made my notes public; the link is in the reply below.
Link to the notes: https://t.co/SyvdgK7S9F
I wrote them to the best of my understanding, but I may well have missed something, so feedback, corrections, and pointers are very welcome.
I received a gold reviewer award from @icmlconf, which provides free registration; I think incentivizing writing good reviews is a smart move. I’m happy to have provided helpful feedback to authors on their papers 😋
LLMs that reason via Chain-of-Thought often keep thinking long after they've already found the answer, wasting tokens and compute. We built Terminator🤖to fix this: a lightweight probe that predicts when an LRM has generated its final answer and exits the CoT early.
My team has been cooking nonstop for a while... and I’m so excited to finally share what we’ve been building!!!
Today, we’re releasing four open models, many of which are the best models of the same size 🥳!!!
tldr;
1) Raon-Speech: 9B SOTA speech LLM
2) Raon-SpeechChat: 9B full duplex model
3) Raon-OpenTTS: 0.3B/1B open-data-open-weight SOTA TTS
4) Raon-VisionEncoder: 0.4B vision encoder trained only with public data
https://t.co/Zjw7QxTJm9
===
1) Raon-Speech (9B)
Raon-Speech is a speech LLM (LLM + speech understanding + speech generation).
It's a bilingual model (English/Korean), and it's ranked #1 on both leaderboards 😎
tldr; it's the best open-model alternative to ChatGPT voice mode.
Model: https://t.co/lxIGpMWaxf
Tech report: https://t.co/byjBUDLhYC
Web demo: https://t.co/J1CWh7LgGO ("Speech Chat" menu here. "auto" is a bit unstable, so use "manual" and choose the language!)
2) Raon-SpeechChat (9B)
While a speech LLM is useful, it’s kind of like a walkie-talkie. A full-duplex model is more like a phone, so it is even more useful in many applications.
That’s why we also built and are releasing Raon-SpeechChat. Again, on several quantitative evaluation metrics, Raon-SpeechChat scored the best on average.
Model: https://t.co/eVcky62mDq
Tech report: https://t.co/byjBUDLhYC
Web demo: https://t.co/J1CWh7LgGO ("Full Duplex" menu here.)
3) Raon-OpenTTS (0.3B, 1B)
We’re also releasing Raon-OpenTTS, a state-of-the-art open-data, open-weight TTS model.
Model + data: https://t.co/WlIgj1OT4R
The 1B model and a detailed tech report are coming soon!
4) Raon-VisionEncoder (0.4B)
Last but not least, we’re releasing Raon-VisionEncoder, a vision encoder trained from scratch using only public data. It closely matchs the SOTA vision encoder quality too!
Model: https://t.co/LO5HxSnjHP
Tech blog: https://t.co/oVzPo4eBMw
===
That’s it!
I’m incredibly proud of what my team has built!
My AI research team at KRAFTON (@Krafton_AI), which undoubtedly is the most cracked team in Korea, has been cooking nonstop for a while for this 😅...
This is just the beginning of our planned model releases, so stay tuned!
ps1/ Ah, by the way, you may ask why “Raon”?
“Raon” is an old Korean word meaning happy.
And, well, we’re kRAftON :-)
ps2/ KRAFTON is one of the four teams participating in Korea’s national frontier-model project, together with SK Telecom. We’re training something very exciting together... and more to come soon!
Want to try it yourself? We're releasing Terminator-Qwen3-8B and Terminator-Qwen3-14B with full vLLM integration. Follow the quickstart on the project page or head straight to our Hugging Face collection
https://t.co/suV5gojfdV
LLMs that reason via Chain-of-Thought often keep thinking long after they've already found the answer, wasting tokens and compute. We built Terminator🤖to fix this: a lightweight probe that predicts when an LRM has generated its final answer and exits the CoT early.
We put together a project page that walks through everything from a high-level overview to the core findings, with figures and a live demo 🚀 If you want the full details, the paper is there too:
Project page: https://t.co/vAvclfFsdn
Paper: https://t.co/SrfFRtIig4
I am excited to announce that our AI institute (Institute for Foundations of Machine Learning, IFML) has been renewed.
IFML was part of the first cohort of AI Institutes announced in 2020. Led by UT Austin, the new award will build on the trajectory of the past five years and develop new foundational tools to advance generative AI. NSF IFML's work on diffusion models is a key technology behind major Google products, powering widely used generative models such as Stable Diffusion 3 and Flux. In it's next phase, NSF IFML will expand generative AI to new domains, including protein engineering, clinical imaging, new methods to handle noisy data, improve agent reliability and open source AI. (1/n)
Thrilled to share that our work received the Outstanding Paper Award at ICML!
I will be giving the oral presentation on Tuesday at 4:15 PM.
@Jaeyeon_Kim_0 and I both will be at the poster session shortly after the oral presentation. Please attend if possible!