Excited to share our latest research and open-source codebase: AgentElect! ๐
We investigate how AI systems can leverage elections to navigate resource-based social dilemmas and drive multi-agent cooperation in common pool resource problems.
๐ Read the paper: https://t.co/4uV0H4FVJ9๐ป
Explore the code: https://t.co/8USGpc62Yv
#MultiAgentSystems #AI #MachineLearning #AIResearch #OpenSource
For everyone visiting Seoul for #ICML2026, my wife @easyminie_ put together a very practical Korea guide based on 20+ years of living here ๐ฐ๐ท
It has a COEX-focused ICML guide, vegan-friendly restaurants, useful apps, cafes, food spots, transport tips, and many curated map lists organized around the conference venue.
Hope it helps people enjoy Seoul during ICML!
https://t.co/heMkSRl6Pg
Multi-domain RLVR trains one model across math, code, logic, science, tables, and simulation. But the key curriculum question is usually fixed or hand-tuned: which domain should we sample next?
๐จ New paper: Transfer-Aware Curriculum (TAC). TAC asks not only โwhere is the model learning right now?โ but also โwill this gradient step help the other domains too?โ
โก The signal is almost free: we reuse GRPO gradients already computed during training, sketch them with random projections, and estimate cross-domain transfer by gradient alignment. No extra rollouts, no held-out probes, <1% wall-clock overhead.
๐ TAC gets the best macro-average accuracy on both Qwen3-1.7B and Llama3.2-3B, with up to +2.8 pts, or 10% relative, over a learnability-only bandit.
๐ฒ Surprise: math was among the least transferable domains, even though RLVR often leans on it most.
Excited for our "Trustworthy AI for Good" (AI4GOOD) Workshop at #ICML2026! As AI agents increasingly affect our lives, it is key to bridge #ResponsibleAI, social good, and governance. Letโs build solutions together!
โฐ Submission deadline: April 30, 2026 (AoE)
๐๏ธConfirmed speakers: @Yoshua_Bengio, Joel Z. Leibo (@jzl86), Maksym Andriushchenko (@maksym_andr), @OanaIgnatRo [More to come!]
๐July 10-11, 2026 ยท Seoul๐ฐ๐ท
๐ https://t.co/2NsL7jFpso
๐ Submit: https://t.co/RcOwxoPRtS
๐ฃ Be a reviewer: https://t.co/8tRAm3onbl
Just went on my first podcast! Enjoyed discussing continual learning for LLM agents and its safety implications with Anna on The Glitchatorio, you can check it out at either of the below links:
Spotify: https://t.co/2ph4bgtk2e
Apple Podcasts: https://t.co/leyRwUyV9W
How can agents understand the world from diverse language? ๐
Excited to introduce Dynalang, an agent that learns to understand language by ๐ข๐๐ ๐๐ฃ๐ ๐ฅ๐ง๐๐๐๐๐ฉ๐๐ค๐ฃ๐จ ๐๐๐ค๐ช๐ฉ ๐ฉ๐๐ ๐๐ช๐ฉ๐ช๐ง๐ with a multimodal world model!
A big day for Python! The steering council has decided to remove the GIL:
- Will unlock fast multithreading
- User code can stay exactly the same
- Experimental support planned for 3.13 (Oct 2024)
https://t.co/XMBca292E3
Wow. @Meta commits to dedicate three engineer-years to implement the removal of the GIL from #Python and fix upcoming compatibility and performance issues with it.
All this dependent on whether the Steering Council accepts PEP 703.
https://t.co/8YqL9KwYL1
STEVE-1 surprisingly also works with visual prompts taken from real life videos, thanks to using MineCLIP's embedding! See๐งต๐
Such a fun project with @Shalev_lif@keirp1@jimmybajimmyba@SheilaMcIlraith
(Credits to @CorruptCarnage and Soil Television for the source videos)
Meet STEVE-1, an instructable generative model for Minecraft. STEVE-1 follows both text and visual instructions and acts on raw pixel inputs with keyboard and mouse controls. Best of all - it only cost $60 to train!
w/ @Shalev_lif@SirrahChan@jimmybajimmyba@SheilaMcIlraith
Excited to share Boosted Prompt Ensembles! Inspired by boosting algorithms, we propose an algorithm that automatically grows an ensembles of prompts to cover a target problem space.
Led by @silviupitis and with @andrew_wang10 and @jimmybajimmyba
One of my favorite results in 2022 was that it's not enough to just think step by step. You must also make sure to get the right answer :D
https://t.co/NbwY5brTgs
(actually a nice insight into a psychology of a GPT; it pays to condition on a high reward)
Neat result from applying automated prompt generation on translation. Referencing "Google Translate" in the prompt improves the performance of the model.
Another interesting prompt found using APE: rather than ask GPT to translate to Spanish directly, ask it to pretend as if it's using Google Translate for stronger performance!
@KabalanFreeman@keirp1@Yongchao_Zhou_@ziwen_h@silviupitis@SirrahChan@jimmybajimmyba We define the score of a prompt x as: E[logP(correct_output | input, x)] =~ 1/n * (logP(correct_output_1 | input_1, x) + โฆ + logP(correct_output_n | input_n, x)) where n is the number of examples in the user provided demo. We then say the prompt with the highest score is best.