AI research agents can generate hundreds of ideas in seconds, but evaluating each can take days of GPU time. When compute is limited, which idea deserves execution?
We introduce AI Research Preference Models (RPMs) to assess ideas & focus compute on the most promising paths 🧵👇
Phenomenal work from the AIRA team.
Huge congratulations to @MartinJosifoski@mahnerak@EdanToledo@RishiHazra95@NickyBaldwin3 for achieving this groundbreaking result.
Watching them work night and day for months on end - this result felt inevitable.
Meta does it best.
As a test of our progress to advance the frontier of AI research, in June we entered the next generation of our autonomous AI research system, AIRA₃, in a live Kaggle competition run by NVIDIA to fine-tune a 30B Nemotron model. The challenge was to teach the model to reason better — all competitors had access to the same information and were graded externally on a private test set.
AIRA₃ placed 8th out of ~4,000 teams to win Gold, outperforming human competitors who had access to the same frontier tools.
We believe this is a reliable signal that AIRA₃ can improve a targeted capability of an AI model at a level similar to human experts.
I quite enjoyed reading Meta's new paper on research preference models (RPMs) this week:
https://t.co/iJgFiaKmKb
It tackles a problem that many people are pondering for automated R&D, namely how does one instil "research taste" in agents as they conduct experiments?
For example, @eliebakouch's scaling of Nano-GPT speed-runs showed that the agents get stuck in local optima where they spend most of the compute budget tuning hyperparameters instead of pushing for novel ideas. Similarly, PostTrainBench shows most models focus on algorithmic tuning instead of curating high-quality datasets (alas, it seems agents hate working with data as much as researchers :))
In the RPM paper, the authors take a different route:
- Treat each experiment as a node in a tree
- Score each node according to e.g. scores on downstream evals
- Mutate nodes to generate candidate experiments
- Use an RPM (aka LLM-judge) to select the most promising node
- Run the actual experiment
- Update tree and iterate
The main thesis is that you can save a lot of wasted experiments by having an RPM guide the search (kind of like how we used PRMs to guide tree search in https://t.co/LT3cERn8Bh)
Although it's unclear whether this tree structure will be needed as agents get better at R&D, one very nice feature is that it can produce a large amount of trajectories that can be used to train domain-specific RPMs which would be quite valuable in hard domains like the natural sciences
A super nice paper and one I recommend reading to better understand how we can make models better at doing research!
Link to the paper: https://t.co/bPtv1huYrt
AI research agents can generate hundreds of ideas in seconds, but evaluating each can take days of GPU time. When compute is limited, which idea deserves execution?
We introduce AI Research Preference Models (RPMs) to assess ideas & focus compute on the most promising paths 🧵👇
A new algorithm for automatic assembly of products is accurate, efficient, and generalizable to a wide range of complex real-world assemblies.
https://t.co/B9QuYTmtIb
This is a letter Feynman wrote to a former student who wrote congratulating him for the Nobel. I’ve posted it before but I really find it worth it to read especially as a student or early stage research person.