We're also releasing MetaGym, a Python library for building meta-environments to support prompt optimisation and self-improvement research for LLM agents.
Find us at the workshops!
https://t.co/mWuzjmEoW5
Our paper on cross-task self-improvement in LLMs is at #ICML2026 (RLxF + AIWILD workshops) this week!
@Mateja_Jamnik@ZifengDing6
How do you train an LLM to learn transferable lessons from its own experience?
Results: reflectors trained on ALFWorld & MiniHack generate better prompts than baselines, generalising to unseen tasks. They improve for 10+ rounds despite 4 rounds of training, with cases of skills transferring across benchmarks.
📄 New paper: "A Minimum Description Length Approach to Regularization in Neural Networks" with Orr Well, Emmanuel Chemla, @roni_katzir, and @nurikolan .
We explore why neural networks often struggle with simple structured tasks.
Spoiler: our regularizers might be the problem.
🧵
📢Paper release📢
What computation is the Transformer performing in the layers after the top-1 becomes fixed (a so called "saturation event")? We show that the next highest-ranked tokens also undergo saturation *in order* of their ranking.
Preprint: https://t.co/ovwI9JpiAQ
1/4
Exciting news! I'll present my poster at #ACL2024 about unsupervised document structure extraction tomorrow (Aug. 12th) at 12:45 PM 🕒 Come say hi and let's chat over the paper! https://t.co/FrZeOlIiuL More details below ⬇️
w/ @GabiStanovsky@yoavgo@allen_ai@nlphuji
The dynamic we capture, where the intermediate answers seem to be significant in the reasoning process, offers a novel cognitive approach for modeling together association and explicit reasoning. This also offers fresh insights into what "thinking" means in the context of AI.
🧠🤖 How do LLMs think? What kind of thought processes can emerge from artificial intelligence? Our latest paper about multi-hop reasoning tasks reveals some new interesting insights. Check out this thread for more details! https://t.co/oiws4IW27e @GoldsteinYAriel@amir_feder
We modeled this transition using a linear regression model, which demonstrates the transformation between the two semantic categories. Our findings indicate a strong connection between the intermediate and final answers categories.
🧠🤖 Does learning in the brain inherently require plasticity? In our latest paper we question this assumption, by leveraging insights into how LLMs "learn". Check out this thread for more details! https://t.co/KTaOC0b0P9
w/ @YuvalShalev1@GabiStanovsky@GoldsteinYAriel