Our presentation with @nissymori1 at ICML is ongoing (Poster #111)!
Come find out about how retrying leads to the emergence of stochastic exploration in reinforcement learning!
#ICML2026
At ICML, tomorrow (Thu 14:30) we’re presenting ReMax: a simple idea that leads to stochastic exploration when greedily maximizing rewards over retries.
What I find exciting is that this idea did not stay in one corner of RL. It now connects to LLM post-training, bandits, image diversity, continuous control, and distributional RL. I created a new thread about our body of follow-up works🧵
#ICML2026 @icmlconf
@Matsuo_Lab@PaavoParmas@sotetsuk@Tdash_Koz@t_kitamura14 I’ll be presenting my poster at Hall A #111 on Wed, July 8, 2026, 10:30 PM–12:15 AM PDT.
Please stop by if you’re around — I’d be happy to chat!
Paper: https://t.co/jtlQbXZfON
Code: https://t.co/M0IR1eKWwk
See you in Seoul!! 🇰🇷
Our paper "Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying" is accepted at #ICML2026 🇰🇷
Proposed a novel exploration objective called ReMax, evaluating best of multiple trials under uncertainty.
The objective comes from the basic question,
Why do RL agents need to explore?
We argue it is because
♻️ Agents are allowed to retry (otherwise, the rational choice is the current best action).
📈 Return is uncertain (otherwise, no point in trying suboptimal actions.)
ReMax naturally captures these intuitions by modeling the distribution of returns and evaluating the maximum over multiple retries, thereby encouraging agents to select actions that are currently suboptimal but highly uncertain.
The diagram is inspired by the Vector Policy Optimization (VPO) paper.
🧵1/n
Introducing Claude Science, a new app designed with every stage of research in mind.
Artifacts traced to their code, environments managed on demand, and 60+ optional scientific databases that you can connect.
Available now in beta.
📢 Call for Papers @eccvconf
The LIMIT Workshop at #ECCV2026 is accepting submissions!! 🚀
How can we learn strong representations when resources are limited? 🤔
We invite submissions on representation learning when data, labels, compute, or other resources are scarce.
⏰ Deadline: July 6, 2026, 23:59 AoE
🔗 https://t.co/hRM8BshLUl
Please consider submitting your work! 🙌
New preprint on statistically principled cheating detection of the Coding agent 🧑💻
The key idea is capped evaluation: create multiple equally valid outputs for a coding task, but evaluate against one randomly sampled choice. With two choices, e.g., 🍎 or 🍊, non-cheating performance is capped at 50% in expectation. Significant cap violations provide a statistical signal of possible test-gaming or leakage.
Led by @skydddoogg — happy to be a co-author!
I want to offer some unsolicited advice to computer vision researchers jumping into robotics. Don't focus too much on VLMs, VLAs etc. That's fine, but the real action is at the sensorimotor level. Most of the open problems in robotics are in manipulation, which is about hand-object interaction, and contacts and forces are central. Proprioception and tactile sensing are as important as vision. Don't get seduced by cherry-picked demos. You can't do robotics without doing robotics.
I’ve been capturing 3D human motion for 30 years and today is maybe the biggest day in that history. We are presenting MAMMA at CVPR (oral session 2A). MAMMA is a markerless multi-camera system that has accuracy similar to marker-based systems.
🎉Our paper will be presented at #CVPR2026 Findings!
"Group-DINOmics: Incorporating People Dynamics into DINO for Self-supervised Group Activity Feature Learning" by Ryuki Tezuka, Chihiro Nakatani @china64681791 , and Norimichi Ukita.
🦖Project: https://t.co/c0b06K8zwT
#CVPR
🎉 Thrilled to share that our paper received the Best Paper Award at the CV4AEC at CVPR 2026!
Huge thanks to my co-authors, the organizers, and the entire CVPR community! 🙏
#CVPR2026#CV4AEC
#ICRA2026 にて主著含む3件,
#CVPR2026 にて1件の発表を行います!
ICRA in Vienna 🇦🇹
- [TuI1I.403] Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
🔗 https://t.co/zO1KH4Nv92
6月5日、Workshop「Geometry in the Age of Data-Driven Robotics(#ICRA2026)」において、#未来創生センター 金井が、当センター最新のカメラ自己校正に関する研究成果を発表します。学習ベースのカメラ自己較正手法の脆弱性を再検討し、提案手法 ZeCNOBA によって対処できる可能性を示します。