That wraps up my time at #icml2026!
Thank you to everyone who stopped by my posters for a chat and I had a great time. This was my first time presenting my work at a conference, and getting to do that in South Korea made the whole experience even more special.
The workshop talks were also really enjoyable and easy to follow. I especially liked a lot of the talks at FAGEN and Agents in the Wild.
A big thank you as well to all the professors and researchers who spared a few minutes to hear me out and share their thoughts.
Iβll definitely miss π°π· but Iβll be back and hopefully next time with more than two days of free time for sightseeing π
Until next time!
@iarthsingh I haven't seen DeepSeek on the web give me any refusals when I try working with bio data but I don't work on bio harm so dunno if that would still work
RL needs more long-horizon tasks beyond math and coding. @niloofar_mire's team at CMU built an environment for drug design.
SMDD-Bench comprises 502 small-molecule design tasks with RDKit, ADMET-AI and Boltz-2 in the loop. The challenges go beyond chemistry: long-horizon planning, exploration, and learning from imperfect feedback are also open problems for RL/ML!
SMDD-Bench is available in our Environments Hub, ready to train with prime-rl. Thank you for sharing with the community!
RL needs more long-horizon tasks beyond math and coding. @niloofar_mire's team at CMU built an environment for drug design.
SMDD-Bench comprises 502 small-molecule design tasks with RDKit, ADMET-AI and Boltz-2 in the loop. The challenges go beyond chemistry: long-horizon planning, exploration, and learning from imperfect feedback are also open problems for RL/ML!
SMDD-Bench is available in our Environments Hub, ready to train with prime-rl. Thank you for sharing with the community!
New blog: Is human taste overrated in harness engineering?
An agent harness is the system around a model. It decides which tools the model can call, what evidence it sees, and what it remembers between steps. Most harnesses are still designed by human taste. Run the agent, read where it failed, change the harness, try again.
That loop is increasingly being automated, with a strong model rewriting the harness itself. So we asked how far that can go. We ran Qwen on three drug design tasks from SMDD-Bench and compared a manually redesigned harness against one produced by automated harness search with Claude in the loop.
No single approach won everywhere. The manual harness pulled clearly ahead when the fix was changing what the agent sees and remembers. Automated search did better once that was in place and the remaining problem was how to search the chemistry itself.
The hard part was rarely implementing the fix. It was diagnosing what kind of failure we were looking at, and that is where human taste still mattered.
Work with @KevinH1119568 , @aviral_kumar2 and @niloofar_mire
Read more π
Students sometimes ask me if it still makes sense, in this accelerating age, to pursue a PhD in AI.
Perhaps counterintuitively, I think it's a great time to do so.
I wrote up some thoughts on this here: https://t.co/RRqGR0qi1R
WE GOT IN!!
#NeurIPS2026 Acceptance!
independent co-first author run.
@SureshRaghu07 absolute pleasure.
congrats everyone who got in
sydney here we comeπ¦
for the past few months i've been asking our models to paint. opus 5.5 is very skilled at emulating different styles
every image here is a python program generated pixel by pixel. there is no image model, and no off-the-shelf art software. instead, it's about 7,500 lines of code using standard libraries to emulate different brush styles. the agents don't use any pictures as reference, instead working only from what they know about each painter
It kinda doesn't matter if the neurips results is before or after ICLR abstract submission since whoever gets accepted for neurips would drop out of ICLR, if they get rejected then they would port the submission and submit to ICLR anyway. Only way to fix this would have been to have the ICLR paper deadline (not abstract) be before Neurips results. The only way to stop the waterfall of papers