Presenting our work which utilizes LLM to assist in individual contribution determination in multi-agent systems! Welcome to my poster if you’re at NeurIPS🌊!
The paper link https://t.co/N6LtQaucfx
🔍Our key contributions💡
- The scale of this reward function naturally reflects an LLM’s ranking uncertainty.
- The method issues uninformative rewards when the LLM is uncertain, avoiding misleading rewards and enabling efficient training even with small error-prone LLMs.
🚀 Excited to share our latest work: extending RLAIF to work well with small language models which may produce incorrect rankings! 🚀 Catch our poster session at #EMNLP in Miami next Thursday!
📄 Paper: https://t.co/eGXXKfVE3p
🎥 Presentation: https://t.co/6amDzfvRFb
Presenting our work which improves robustness of RL based on noisy LLM feedback today at @RLBRew_2024 ! If you are at RLC and would like to chat, please reach out :)!
The paper link is here 👉https://t.co/my2sRztnQe