Knew we were due for a turnaround!
Our @SorareNBA recommended optimal lineups smashed in GW28. Even with 0% xp and collection boosts, it would have taken down 🥇 first place in Superrare Champion this GW.
My intuition to explain GPT-4 bias towards LLMs is:
• InstructGPT/GPT-4/ChatGPT are rewarded (through RLHF) for producing responses that are judged ‘desirable’ by human evaluators
Our LLM leaderboard broke the internet💥 and took the community by storm🌀
We have now expanded our leaderboard to include Human evals in partnership with @scale_AI and GPT4 evals🚀
The most fascinating result to me is how the most human-aligned model, GPT4, is actually not that well-aligned with human preferences🤯
Blog: https://t.co/kN09ajOeWE
Leaderboard: https://t.co/PbUNObr7yE
• Models such as Vicuna or Alpaca are fine-tuned with data generated from InstructGPT/GPT-4/ChatGPT models, they learn how to generate ‘desirable’ responses, but they also learn to reproduce the underlying GPT-preferred “style” of writing
@nazneenrajani@scale_AI Thanks for sharing! I think this is crucial work. GPT-4's preference for LLM content could have huge (and bad) consequences. We can imagine LLM cover letters outranking superior human ones in automated systems. A new bias is being created between human and LLM-augmented humans
Our LLM leaderboard broke the internet💥 and took the community by storm🌀
We have now expanded our leaderboard to include Human evals in partnership with @scale_AI and GPT4 evals🚀
The most fascinating result to me is how the most human-aligned model, GPT4, is actually not that well-aligned with human preferences🤯
Blog: https://t.co/kN09ajOeWE
Leaderboard: https://t.co/PbUNObr7yE
🏀#SorareNBA With the end of the playoffs fast approaching, I’m planning to share a Sorare NBA 2022-23 Season data recap with more details than the official one. Are there any specific insights or stats you’d like to see in this recap?
Open-source LLMs are charging ahead in capabilities, but there's an incentive mismatch around building harmless models too:
* why the harmlessness gap is emerging
* red-teaming scorecards v leaderboards
* missing pieces for constitutional AI
https://t.co/WSkBSL2jMl
"instead of alarming the public with ambiguous projections about the future of AI, we should focus less on what we should worry about, and more on what we should do" https://t.co/ZZAVzAb1WX excellent thoughts by @sethlazar@jeremyphoward@random_walker#ResponsibleAI#EthicalAI
The orgs most responsible for deploying AI in unsafe ways are also the loudest when it comes to talking about AI safety, and they focus the discourse not on actual AI-linked risks (bias, IP theft, worker disempowerment, deep fakes...), but on imaginary ones (SkyNet). Distraction?
I’ve worked my whole life on AI because I believe in its incredible potential to advance science & medicine, and improve billions of people's lives. But as with any transformative technology we should apply the precautionary principle, and build & deploy it with exceptional care
yay the ability to share ChatGPT conversations is now rolling out. I can share a few favorites.
E.g. GPT-4 is great at generating Anki flash cards, helping you to memorize any document. Example:
https://t.co/TVmTlXxbDB
Easy to then import in Anki: https://t.co/kaHdvO2FFc
New post: the AI Canon
We share all the papers, posts, articles, courses, and videos we've relied on to get smarter about LLMs and modern AI
Compiled by @derrickharris@appenz and myself
https://t.co/bZ5tht1xWE
With more powerful AI systems comes more responsibility to identify novel capabilities in models. 🔍
Our new research looks at evaluating future 𝘦𝘹𝘵𝘳𝘦𝘮𝘦 risks, which may cause harm through misuse or misalignment.
Here’s a snapshot of the work. 🧵 https://t.co/Y499hpV4no
So excited to share what we've been working on in the past couple of weeks!
If you're new to Transformers, you can now just *ask* the agent what you want to achieve. In plain English. It will select the right tool (and model) for you.