Hello world :)
We are BOLD — the British Open-ended Learning and Discovery Lab!
BOLD is a new academic research lab fully focussed on paradigm breaking discoveries in fundamental AI. We work towards more efficient & open AI that is built around human needs and capabilities.
To pursue these breakthroughs, we pioneer new modes of collaboration in academia that are more focussed, resourced, agile, and collaborative. Rather than fragmenting resources, today we are sunsetting 5 of the UKs leading AI labs to join forces under our joined scientific vision.
Our vision is centered around three pillars:
⚡ Beyond backpropagation – questioning the foundations of the field.
🤝 Human-centric learning & discovery – treating humans as core to our algorithms
🤖 Embodied learning – fast learning and adapting methods that deal with the messy real world
BOLD is backed by @UKRI_News and @EPSRC with £30M – and this is just the beginning. We are urgently looking for partners and sponsors to 10x this.
👉 https://t.co/eFVFW31mqz
👉 https://t.co/Eoad4G18KL
@j_foerst, @CULLYAntoine, @tonizza82, @shimon8282, @tonizza82, Ani Calinescu & @_rockt
Excited to see our Centaur project out in @Nature.
TL;DR: Centaur is a computational model that predicts and simulates human behavior for any experiment described in natural language.
AI Meets Game Theory:
New study reveals that while today’s AI is smart, it still has much to learn about social intelligence.
Read more:🔗 https://t.co/4jtNtt8iqb
#AI#GPT4#GameTheory#SocialAI
Our paper is now out in Nature Human Behaviour! 🎉 We use games from behavioural economics to explore how LLMs behave in repeated social interactions, revealing both self-interested strengths and coordination blind spots, and propose strategies to improve AI-human collaboration.
What happens when Large Language models play repeated economic games? This new study from @elifakata et al investigates across a range of models and games. @cpilab
https://t.co/ID2d6mg5sz
What happens when Large Language models play repeated economic games? This new study from @elifakata et al investigates across a range of models and games. @cpilab
https://t.co/ID2d6mg5sz
🚨 We're hiring! If you're excited about 🤖 ML/LLMs, 🧠 cognitive science, or 💭 computational psychiatry, come join us in Munich.
Two fully funded PhDs @HelmholtzMunich: great mentorship, international vibe, lots of room to grow.
📅 May 16th
🔗 https://t.co/j0mwz9LFNx
In previous work we found that VLMs fall short of human visual cognition. To make them better, we fine-tuned them on visual cognition tasks. We find that while this improves performance on the fine-tuning task, it does not lead to models that generalize to other related tasks:
Our paper (with @elifakata, @MatthiasBethge, @cpilab) on visual cognition in multimodal large language models is now out in @NatMachIntell. We find that VLMs fall short of human capabilities in intuitive physics, causal reasoning, and intuitive psychology. https://t.co/jGIChi7U05
Excited to announce Centaur -- the first foundation model of human cognition. Centaur can predict and simulate human behavior in any experiment expressible in natural language. You can readily download the model from @huggingface and test it yourself: https://t.co/nLBYhHpCtT
🚨Don't forget to join us tomorrow at Shaw Library (Floor 6, Old Building, LSE) for a full-day event showcasing latest research on generative AI in social science research from scholars at @LSEnews, @UniofOxford, @BYU and @uni_tue.
We look forward to welcoming you all!
🚨Pre-print alert:🚨
Have we built machines that think like people?
In new work, led by @lucaschubu and @elifakata and together with @MatthiasBethge, we assess multi-modal #LLMs reasoning abilities in three core domains: intuitive physics, causality, and intuitive psychology.
For more than two years now, a part of our lab has been focusing on using the behavioral sciences to better understand Large Language Models. Time for a small recap of what we have found. As these agents will interact with us increasingly more often, how do they actually behave?
It’s time for some game theory!
What happens if we ask AI to play common economic games? They do very well at games where being selfish is optimal (Prisoner’s Dilemma) but worse at those that require coordination. GPT-4 did best, especially with prompting https://t.co/JwxDqlbG3P
Playing repeated games with Large Language Models
propose to use behavioral game theory to study LLM's cooperation and coordination behavior. To do so, we let different LLMs (GPT-3, GPT-3.5, and GPT-4) play finitely repeated games with each other and with other, human-like strategies. Our results show that LLMs generally perform well in such tasks and also uncover persistent behavioral signatures. In a large set of two players-two strategies games, we find that LLMs are particularly good at games where valuing their own self-interest pays off, like the iterated Prisoner's Dilemma family. However, they behave sub-optimally in games that require coordination. We, therefore, further focus on two games from these distinct families. In the canonical iterated Prisoner's Dilemma, we find that GPT-4 acts particularly unforgivingly, always defecting after another agent has defected only once. In the Battle of the Sexes, we find that GPT-4 cannot match the behavior of the simple convention to alternate between options. We verify that these behavioral signatures are stable across robustness checks. Finally, we show how GPT-4's behavior can be modified by providing further information about the other player as well as by asking it to predict the other player's actions before making a choice. These results enrich our understanding of LLM's social behavior and pave the way for a behavioral game theory for machines
paper page: https://t.co/72dKsSda7k