Our latest work demonstrates how industry-scale Flywheel systems can iteratively drive model improvements even when the objectives are not easily tractable!
1/5 🤔 LLMs can solve olympiad math and write production code. But can they hold a conversation that's actually fun — one that people want to keep coming back to? 💬✨
We present CharacterFlywheel— an iterative process optimizing LLMs for real human engagement and character steerability, while maintaining rigorous safety protocols 🔒. Tested across Instagram, WhatsApp & Messenger 📱with millions of users — where they can create, share, and chat with their own AI characters 🤖.
📄 paper: https://t.co/WJ9xKz89Dt
https://t.co/9lA3DdkWuM
@natolambert Totally agree! Also sharing our recent paper that shows RL environments with high state information richness & high planning complexity are crucial for generalization: https://t.co/y1DexWPgeU
Our latest work sheds light on what to "scale" when building generalizable agents:
👉 Prefer training envs with 📚 high state information richness (high perception load & info volume) + 🧠high planning complexity (long task horizon & branching factors)
👉 The *complexity structure* matters more than realism: Hard problems in "toy" domains like Sokoban / BlocksWorld can be more useful than easy problems in more realistic domains like ALFWorld.
‼️📉 Be careful about your mid-training datamix and strength: warmup and mid-training help prevent catastrophic forgetting during RL but undermines generalization to domains that are not covered
💡Applying lightweight state randomization/augmentation helps!
📄 Paper: https://t.co/qmaIjMFYwJ
Our latest work sheds light on what to "scale" when building generalizable agents:
👉 Prefer training envs with 📚 high state information richness (high perception load & info volume) + 🧠high planning complexity (long task horizon & branching factors)
👉 The *complexity structure* matters more than realism: Hard problems in "toy" domains like Sokoban / BlocksWorld can be more useful than easy problems in more realistic domains like ALFWorld.
‼️📉 Be careful about your mid-training datamix and strength: warmup and mid-training help prevent catastrophic forgetting during RL but undermines generalization to domains that are not covered
💡Applying lightweight state randomization/augmentation helps!
📄 Paper: https://t.co/qmaIjMFYwJ
🤔 Motivation:
Just like AlphaGo, given enough compute and proper configurations, RL agents can achieve superhuman performance in almost all training tasks ⚠️ However, LLM agents today are being used in tasks far beyond what limited post-training data can cover — a challenge highlighted by many researchers like @ilyasut, @MiniMax_AI (in https://t.co/C4SBFORt1c) and @Kimi_Moonshot in K2.5 tech report.
People are putting lots of efforts into building diverse RL environments. But what kinds of environments are more useful for building **generalist** agents that are able to solve tasks beyond their training tasks?
🚀 Instead of optimizing one single benchmark, we look for drivers of transfer in our latest paper: https://t.co/cpgXXbiday
Joint work with MSL @MetaAI (@ZhihanLiu21628@GuanSuns@EasonNie@KaiZhang_CS@nazzhang) where @ZhihanLiu21628 interned. (1/n)
📢 Yochanite Lin Guan (@GuanSuns; https://t.co/lPinpBEGp1), currently a Research Scientist with @Meta GenAI, will be defending his @SCAI_ASU PhD dissertation tomorrow 10/21 👇. He developed several 🔥 techniques for taming the notorious sample complexity of #RL systems..
While LLMs cannot verify the correctness of agent plans, they can be good at capturing the "style" of desirable behaviors.
Come stop by #COLM2024 poster #34 Wednesday Morning to see how we utilize VLMs as a knowledge source of common human preferences for embodied agents!
This work is also part of an ongoing effort in our lab (supervised by @rao2z ) to identify constructive roles that LLMs/VLMs can serve in planning tasks.
👉 Relevant paper: https://t.co/CQB3fGzYIM
While LLMs cannot verify the correctness of agent behaviors, they can be good at capturing the "style" of desirable behaviors. We set out to find the effectiveness of VLM critics of undesirable behaviors.
Robots now understand decisions via SERLfD – Self-Explanation for Reinforcement Learning from Demonstrations, transcending traditional robot learning from human demos. Check out our article and slides; find more insights at poster 409 during AAAI-24 on 2/23, 7-9 PM!🤖🌐 #AAAI2024
📢📢 Come stop by #AAAI2024 poster spots 644 & 645 (the very last spots) this evening to hear me explain the neat work of
👉@ZahraZ__ & @sailiks on explaining allocations to humans
👉@YantianZha & @GuanSuns on using self explanations to learn from ambiguous demonstrations
[Note that at my request, the posters were placed in contiguous spots instead of their original spots 409/435 so I didn't have to master the art of being in two places at one time.. 😋]
📣Chalk it to mad masochism, but we will be updating and reprising our tutorial on LLMs and Planning at #AAAI2024 Vancouver, BC. (Wed Feb 21st, 2-6PM 👉https://t.co/mmO2FDFRbk. We are told that ~300 folks signed up already..😱). w/ @karthikv792@GuanSuns
📢Don't miss out our poster session tomorrow (Dec 12) starting at 5:15 p.m. CST in the Great Hall & Hall B1+B2 #1525
This is a joint work with @karthikv792@sarath_ssreedh and @rao2z
Paper homepage: https://t.co/uT8KNkuv3y
Presentation at NeurIPS 2023: https://t.co/3o3V5HQCig
Interested in building #LLMAgent but tired of endlessly fixing flawed plans? 😫🤦♀️
Our #NeurIPS paper reveals LLMs aren’t designed for action sequencing! Instead, they can be useful for extracting planning knowledge and generating codified models that drive external planners!💡1/
Having missed attending #ICLR2023 in-person, we are also re-presenting the paper at the #ICML2023 ILHF workshop on Saturday, July 29th. Feel free to drop by if you are on Waikiki ... 🏖️ Mahalo 🙏 2/
📢 Our #ICLR2023 paper on learning and leveraging relative behavioral attributes to tame the sample complexity of #RLHF (joint with @karthikv792 & @rao2z ) is now featured on @mtlaiethics ! 1/
👉 https://t.co/0dJXAAENw0