Working on enhancing reasoning capabilities of LLMs | PhD Student @UniBasel under @ilijabogunovic and @AurelienLucchi | ex-@EPFL ๐จ๐ญ and @unict_it ๐ฎ๐น
๐ฏ Learning is most effective at the frontier of capability: problems that are too easy or too hard teach nothing.
For LLM reasoners trained with GRPO this is literal: problems the model always or never solves give zero gradient.
We introduce Frontier Learning๐๐งต
โ Takeaway: good training data isn't something you pick once before training. It has to keep moving with the model, toward what it can almost do. Frontier learning does exactly that, and the longer you train, the more it pays
DiffusionGemma is one of the first large, open-weight uniform diffusion LLMs.
But how can the community actually post-train it?
New blog ๐: what works, what breaks, and the SFT objective that wins. ๐งต
Weโre hiring! ๐จ๐ญ
Join the University of Basel as a PhD student or postdoctoral researcher in the field of LLM training.
Work on:
๐ค LLM pre- and post-training
๐ฏ Reinforcement learning and reasoning
๐ Reliable generalization
โ๏ธ Large-scale experiments
๐ Basel, Switzerland
Heading to ICML? Check out this RL work from my student @xiaohang_tang and collaborators on teaching diffusion language models to reason by denoising sequences. Featuring wd1, wd1++ & GDSD
Hiring 2 summer ML research interns at the University of Basel ๐จ๐ญ.
Research topics: RL/diffusion LLM post-training, reasoning, or LLM orchestration. Possible fully funded PhD offers to follow.
I'll be at ICLR this week and happy to chat.
Apply: https://t.co/gOdLRv771L