I am pleased to announce another update to my RL tutorial (https://t.co/fnHAmyg1mH). This time I have added code for RLFT for multi-turn LLM agents, using the awesome Tinker library from @thinkymachines, and the simple ReBN training loop from GEM by @zzlccc et al. With ~100 lines of simple python running on your laptop, you can train an agent based on Qwen3-4B-Instruct to play "guess the number" in 20 minutes.
🎓Getting started in information geometry in 1 hour!
- 10-page pdf short introduction (AMS Notices):
👉https://t.co/6hRmU1qR7r
- 45 min. video tutorial:
👉 https://t.co/ONvjDsH8Ag
I've greatly expanded my chapter on Bayesian clinical trial design with examples of Bayesian power and sample size simulations for time-to-event and ordinal outcomes, incorporating uncertainty in effect size to detect ... https://t.co/jDPYp3i8Gc @vandy_biostat
Another AI paradox: people are excited about LLMs, some even think that AGI is just around the corner. But some students are depressed how they can still get a PhD. Is it becoming pointless?
Some personal notes on this. (1/8)
New book draft from @BachFrancis
"Learning Theory from First Principles"
https://t.co/Hg5EY0PmE2
Lots of great material, quite up to date on latest theory.
Also, we recently posted this work https://t.co/bOWAF0UmBr on the theory of policy gradients for reinforcement learning! A long time in the works, this paper finally gets a handle on function approximation with policy gradient methods.
For folks looking for a thorough intro to the mathematical foundations of reinforcement learning: Video lectures for Bertsekas’ course on RL and control are now available here: https://t.co/cJrYJqX1Rn
Is your estimator of the average treatment effect simply the best, better than all the rest? Now is your chance to prove it once and for all (or at least until next year). The annual Causal Inference Data Challenge is now live! Datasets available at https://t.co/XO5JSpUibF
blockCV: an R package for generating spatially or environmentally separated folds for k‐fold cross‐validation of species distribution models - Valavi - - Methods in Ecology and Evolution - Wiley Online Library https://t.co/uvNPwVsp7R
Constantinos Daskalakis has been awarded the 2018 Rolf Nevanlinna Prize, one of the highest honors in theoretical computer science, for his explorations of machine learning and core questions in economics. #icm2018#FieldsMedal2018 https://t.co/0M25htE8Ga
do explanations help in recommender systems, and if so, how can they work with bandits? recent slides from a talk @SpotifyResearch during the Spotify ML day in Stockholm https://t.co/81qsSM0k2b