Excited to share our new paper on implicit safety alignment from crowd feedback that was accepted at ICML! We learn shared safety constraints from crowd feedback and show this can help alleviate reward mispecification. Kudos to @QianLin44 for this really fun piece of work!
What if your task reward is imperfect β especially when it misses safety-related objectives? Maybe crowd preferences can help.
Excited to share our paper, Implicit Safety Alignment from Crowd Preferences, accepted to ICML 2026! Grateful to @daniel_s_brown for his guidance and support on this work.
Paper: https://t.co/UbBhas54Kw
π§΅1/7
This motivates per-user adaptation and personalization, an exciting area of future work! This has been a really fun collaboration within the @URoboticsCenter. Kudos to co-Authors: Colin Rubow, Eric Brewer, Ian Bales, and Haohan Zhang!
https://t.co/XaJV2TJSbZ
π§΅8/8
Our paper "A Multi-Layer Sim-to-Real Framework for Gaze-Driven Assistive Neck Exoskeletons" is being presented this week at ICRA 2026!
https://t.co/XaJV2TJSbZ
π§΅1/8
We find that our two of our three data-driven gaze controllers are competitive or superior to prior state-of-the-art. Interestingly, there is no single overall "best" controller. We found three non-dominated controllers and which one is best depends on the individual user.
π§΅7/8
This work offers a new mechanistic lens for understanding why normalization, resets, and classification-based value learning work in deep RL β and provides practical guidance for choosing among them depending on reward sparsity.
Full paper: https://t.co/Tm76XFMSmm
7/N
Super excited about our lab's new paper led by @zifan_w: "Understanding the Effects of Neuron Dominance in Deep Reinforcement Learning", that has been published in Transactions on Machine Learning Research (TMLR)!
https://t.co/Tm76XFMSmm
1/N
Excited to share that our paper, "Understanding the Effects of Neuron Dominance in Deep Reinforcement Learning", has been published in Transactions on Machine Learning Research!
Full paper: https://t.co/M6U7X3bKjC
Work with Qian Lin, Blake Lawlor, Haijun Zhao, and Daniel Brown.
4. We provide a novel theoretical explanation for why classification losses like HL-Gauss resist representation collapse: unlike MSE regression, the softmax nonlinearity prevents gradients from vanishing even when the representation collapses to zero.
6/N
How can you do imitation learning for multi-agent systems if the demonstrator is just a single human? Unless you're a doctor octopus genius, this is really hard! We study this problem in a new paper by my student Connor (@connormat) that will be at ICRA 2026!
The Kahlert School of Computing at the University of Utah is hiring for multiple faculty positions! We're especially interested in growing in the areas of human-centered AI and robotics! https://t.co/WZfJnJ5PId
We hope this work can help inspire the development of better AI alignment tests and evaluations for LLM reward models.
Check out the workshop paper here: https://t.co/MWTPYSSTGS
8/8
Can you trust your reward model alignment scores?
New work presented today at the COLM Workshop on Socially Responsible Language Modelling Research by Purbid Bambroo in collaboration with @anmarasovic that probes LLM preference test sets for redundancy and inflated scores.
1/8
We applied this approach to RewardBench and found evidence that much of the data in safety and reasoning datasets may be redundant (44% for safety and 24% for reasoning) and that this can lead to inflated alignment scores.
7/8