Blog: Towards Proactive Value Alignment
How can language models proactively query to align with individual user preferences? We frame this as a reward-uncertain MDP with Expected Value of Information as the objective.
https://t.co/EBv9gIcP4m
#AI#RLHF#LLM#ValueAlignment
It is just so sad that the #NeurIPS2024 main conference ended with such a racist remark by a faculty when talking about ethics. How ironic!
I also want to commend the Chinese student who spoke up right on spot. She was respectful, decent, and courageous. Her response was exemplary: she began by acknowledging the speaker’s efforts, then gave the speaker an opportunity to clarify (though, regrettably, the speaker’s reply only reinforced her bias), and ultimately called attention to the inappropriate racial bias and offered constructive suggestions. Thank you for speaking out!
Combining LLMs with formal math (e.g., theorem proving & autoformalization) is a promising avenue for advancing AI4Math. This long overdue survey provides comprehensive pointers for anyone wanting to dig deeper into this field. Great work by @_ZhaoyuLi, Jack, Logan, Qidong, Zenan, Xian, and @XujieSi!
arXiv: https://t.co/7oSvozsIDi
GitHub: https://t.co/Y25ILnj6Ii
The LLM Alignment team at IBM Research is looking for a talented PhD student for a summer internship at the MIT-IBM Watson AI Lab in Cambridge. If you are working on topics aimed at enhancing model performance and safety, please reach out.
The Lunar Lab at @GeorgiaTech is actively looking for multiple highly motivated Ph.D. students to join us in Fall 2024. We are interested in robot perception, robot learning, and autonomous navigation. Please repost to share widely!
Detailed information: https://t.co/VcHpsAR1e9
Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
Overviews techniques to understand, improve, and complement RLHF in practice
https://t.co/pBX2mAqd1C
How to harness foundation models for *generalization in the wild* in robot manipulation?
Introducing VoxPoser: use LLM+VLM to label affordances and constraints directly in 3D perceptual space for zero-shot robot manipulation in the real world!
🌐 https://t.co/FhBazTzi7Z
🧵👇
Can LLMs generate mathematical proofs that can be rigorously checked?
We release LeanDojo (https://t.co/zkOyW4FoDx): an open-source playground consisting of toolkits, benchmarks, and models for LLMs to prove formal theorems in the Lean proof assistant.
Key features:
- Tools for data extraction and interacting with Lean.
- Fine-grained annotations of premises (e.g., existing lemmas) in proofs: where they are used and defined.
- LeanDojo Benchmark: 97K human-written theorems/proofs for developing machine learning models on theorem proving.
- ReProver (Retrieval-Augmented Prover): the first LLM-based prover augmented with retrieval for premise selection.
We open-source everything, providing the first set of open-source LLM-based theorem provers without any proprietary data, model, or code.
Can LLMs reliably solve long-horizon planning problems?
LLM+P: Empowering Large Language Models with Optimal Planning Proficiency
Link: https://t.co/wNmIQqXQtC
Code: https://t.co/efVlbv67Df
@alexbalandin@minimario1729@_akhaliq@huggingface Our tree search algorithm will decide when to generate codes that are similar to known-to-be-good codes (exploitation), and when to generate codes that are dissimilar to previous codes (exploration).
@alexbalandin@minimario1729@_akhaliq@huggingface First author here :) Good question! I believe contrastive search aims to generate more diverse samples. Our work instead takes the quality (i.e. pass rates on the test cases) of the generated codes into consideration...
Planning with Large Language Models for Code Generation
empirically evaluate framework with several large language models as backbones on public coding
challenge benchmarks, showing that 1) it can generate programs that consistently achieve higher performance compared with competing baseline methods; 2) it enables controllable code generation, such as concise codes and highly-commented codes by optimizing modified objective
abs: https://t.co/clrqPv2uu9
project page: https://t.co/9kAPPHqLKN
github: https://t.co/dspkOzUrj4
#ICML22 Can we build autonomous agents to leverage prior experience and learn novel tasks from a handful of demonstrations as we humans? Please check our Prompt-DT, which leverages the sequential modeling ability of the Transformer architecture to achieve few-shot adaptation!
Prompting Decision Transformer for Few-Shot Policy Generalization
Prompt-DT is a strong few-shot learner w/o any extra finetuning on unseen target tasks.
proj: https://t.co/mw7wJWMmkU
abs: https://t.co/66aLhY061l
Prompting Decision Transformer for Few-Shot Policy Generalization
abs: https://t.co/bD2f4SjRP6
project page: https://t.co/bfTku3MlVw
experiments in five MuJoCo control benchmarks show that Prompt-DT is a strong few-shot learner without any extra finetuning on unseen target tasks