“What objective function do we want AI to optimize for? If we aggregate values from society, what weights do we use, and whose values?”
Learn more about SRI Grad Affiliate @silviupitis's research, supported by an @OpenAI Superalignment Fast Grant.
🔗 https://t.co/Q4mSD6pC5g
📝 How do you choose which language model to use? Quantitative benchmarks can be uninformative and fall prey to Goodhart's Law, and even Chatbot Arena performance can be optimized for.
In our new preprint, we propose generating qualitative report cards... 🧵
🔍 Current LLM evaluations fall short:
• Lack nuanced understanding of model capabilities
• Overly focused on quantitative metrics
• Difficult for humans to interpret
Introducing LLM Report Cards! A novel approach for qualitative, interpretable model evaluation.
1/N
@denny_zhou Neither answer is very good.
A better response would first seek a common scenario where your assumption (11.11 > 11.8) is correct... for example, as a package version. Python 3.11 is indeed "larger" than Python 3.8.
By fine-tuning a 7B parameter reward model on RPR, we achieve better context specific performance than unconditioned models, including larger models used in llm-as-a-judge mode.
See the paper for more details: https://t.co/6OTqd8BXjZ
With @ZiangXiao, Nicolas Le Roux & @murefil
When evaluating LLMs, added context such as a criteria or user profile, may be critical for determining preferred behavior.
But can reward models effectively incorporate this additional context?
📝 New paper: https://t.co/6OTqd8BXjZ
🤗 Dataset: https://t.co/QrKZ6rZMdK
However, we noticed that current models may fail to consider context.
So we synthesized a Reasonable Preference Reversal (RPR) dataset, where every preference query comes with an alternative context under which preference reverses.
@DavidMSidhu Use dropbox or other auto syncing drive for storage. Then use vscode with Foam plugin.
Better / more flexible than obsidian IMO.
This is for personal notes. If you intend to share, Notion / Google docs.
We are presenting ToolEmu at #ICLR2024 tomorrow!
⏲️ Friday 4:30pm-6:30pm CEST
📍 Spotlight poster session, Hall B #80
I won't be able to attend ICLR this year but don't miss the chance to meet our amazing collaborators!
Here's what I see as a likely AGI trajectory over the next decade.
I claim that later parts of the path present the biggest alignment risks/challenges. The alignment world has been focusing a lot on the lower left corner lately, which I'm worried is somewhat of a Maginot line.
ToolEmu has been accepted at #ICLR2024 as a Spotlight presentation🔥
Explore our LLM-based emulation framework for identifying LLM agent risks at scale!
🎯 Demo: https://t.co/o5PdJhgXw2
📄 Paper: https://t.co/pfHb81PyLb
🔗 Code: https://t.co/yJd2pgNNPe
🧵⬇️
@ylecun 1. Some control does not mean complete control.
2. If some of my fellow humans had access to weapons of mass destruction ... damn right I'd be scared of them.
3. Institutions do the wrong thing all the time.
Our team at MSR Montréal is looking for interns!
Subjects range from efficient modular adaptation to building complex systems by stacking LLMs.
Consider applying here : https://t.co/dhDlciSKLL
I will be at #NeurIPS2023 Dec 11-16
Shoot me an email to connect! Particularly interested in:
- LM eval for long-horizon / agents
- Alignment / rewards generally
Will present my paper on multi-objective reward aggregation at Poster sess 6 Thurs eve
(https://t.co/ZSZxbrlDQl)
#OpenAI’s GPTs & Assistants APIs are a blast, making it much easier to build customized agents with new tools. But are they safe to deploy? 🚨
A simple & quick test against prompt injections reveals that it is fairly easy to make GPTs delete all your files 💀
OpenAI just announced GPTs and the Assistants API for “ helping developers build agent-like experiences”, but what does that mean and how does it change how we should govern AI?
Some early thoughts relating to my ongoing work 🧵:
When given context about a “green square” and a “blue circle”, how do language models bind corresponding shapes and colors?
Using causal experiments, we find that large enough language models learn simple structured representations for binding!
A thread (1/n)
RLHF typically assumes that all training feedback comes from a single teacher, but teachers can disagree up to 37% of the time in practice. In our new paper, we introduce active teacher selection to learn from different teachers. (1/n)