Action chunking is drawing growing interest in RL, yet its theoretical properties are still understudied.
We are excited to share some insights on when we should use action chunking in Q-learning + a new algo (DQC) to tackle hard long-horizon tasks!https://t.co/izVWQBgH3c🧵1/N
Excited to release Holosoma, our new full-stack, open-source humanoid learning library from Amazon FAR!
This has been a huge deal for me -- It enables an incredibly fast research cycle. I'm now able to train and deploy a new policy on real hardware in just 20 minutes with an RTX 4090.
You can simply train RL policies, benchmark in Isaacgym, Isaacsim, MuJoCo, and run sim2real deployment with just a few commands in a single codebase!
Often distracted by shorts, videos, and digital noise? Introducing INA, an AI agent for intentional living. State your intention — the agent guides you back when you drift off-task. Paper:
https://t.co/qDR05lRLgw
Website & code (Mac app): https://t.co/3QDiyKcBsh
Thread below 1/N
Excited to introduce flow Q-learning (FQL)!
Flow Q-learning is a *simple* and scalable data-driven RL method that trains an expressive policy with flow matching.
Paper: https://t.co/kjaeqHcBFh
Project page: https://t.co/D8vFcZib1F
Thread ↓
Excited to share Adaptive Low-Pass Guidance (ALG): a simple training-free, drop-in fix that brings dynamic motion back to Image-to-Video models! Demo videos, paper, & code below!
https://t.co/4NzYDfCFSb (🧵 1/7)
Excited to present FastTD3: a simple, fast, and capable off-policy RL algorithm for humanoid control -- with an open-source code to run your own humanoid RL experiments in no time!
Thread below 🧵
"Balancing helpfulness and safety is a real challenge for AI agents in mobile environments."
@kimin_le2 from @kaist_ai introduces MobileSafetyBench—a tool for testing AI safety in mobile devices.
At #NeurIPS to present papers! Details in 🧵.
I’ll be in Vancouver (Dec 10-15) and would love to connect - love to discuss topics like efficient LLMs, agents, and LLM safety.
Also looking for PostDoc and RS positions. Happy to connect if you’re hiring or know of relevant roles!
D4RL is a great benchmark, but is saturated.
Introducing OGBench, a new benchmark for offline goal-conditioned RL and offline RL!
Tasks include HumanoidMaze, Puzzle, Drawing, and more 🙂
Project page: https://t.co/RD68e07ds9
GitHub: https://t.co/22gGfaXPIJ
🧵↓
Latest work on leveraging prior trajectory data with *no* reward label to accelerate online RL exploration!
Our method leverages our prior work (ExPLORe) and skill pretraining to achieve better sample efficiency on a range of spare-reward tasks than all prior approaches!
🚀 First step to unlocking Generalist Robots! Introducing 🤖LAPA🤖, a new SOTA open-sourced 7B VLA pretrained without using action labels.
💪SOTA VLA trained with Open X (outperforming OpenVLA on cross and multi embodiment)
😯LAPA enables learning from human videos, unlocking potential for robotic foundation model
❗Over 30x pretraining efficiency for VLA training
🤗Code and checkpoints are all open-sourced!
Is "offline RL" in offline-to-online RL really necessary?
Surprisingly, we find that replacing offline RL with *unsupervised* offline RL often leads to better online fine-tuning performance -- even for the *same* task!
Paper: https://t.co/1jW2pgOvtB
🧵↓
We show that this improves performance across diverse benchmarks like AntMaze, Kitchen, ExORL, and Adroit.
Check out our paper for more information!
Paper: https://t.co/1jW2pgOvtB
co-led by @JunsuKim97 (during his visit at Berkeley), w/ @svlevine
Excited to present RSP: Representation learning with Stochastic frame Prediction, a new method that learns image representations from videos by training stochastic frame prediction model 🖼️
#ICML2024
Paper: https://t.co/FhMZxsIxnr
🤔 How can we detect texts generated from recent powerful LLMs such as GPT4 and Llama3?
🕵 Use your reward model!
Arxiv: https://t.co/uYwOm10UpJ
Project page: https://t.co/F1zdtsqgwA (1/N)
METRA is the *first* unsupervised RL method that can learn diverse locomotion skills purely from pixels, and is one of my favorite works!
METRA got accepted to ICLR 2024 (Oral), and come to the sessions this Wednesday!
Oral: Wed 4p, Halle A 2
Poster: Wed 4:30-6:30, Halle B #161
🤩Realistic benchmark --> practical AI agents
Excited to share "B-MoCA", our new benchmark for evaluating mobile device control agents across diverse device configurations📱
Arxiv: https://t.co/FbGbdiO1ME
Website (with code): https://t.co/ocYFmJVtgA
🧵1/N
Try on different clothes and explore new styles without making a purchase.
IDM-VTON, our diffusion based try-on model lets you wear clothes virtually.
Visit our project page for more examples and try your own photo. 👇
https://t.co/JYlaeDBXTn