OpenRSI-X feels like an important step in that direction. Excited to contribute, and looking forward to seeing more researchers and domain experts join us!
What excites me most is moving beyond “better models” toward agents that can turn ideas into experiments, test them rigorously, learn from failure, and eventually propose experiments that scientists would actually want to run.
As AI systems enter a recursive self-improvement loop, the central question is whether it can systematically move beyond human-designed methods to genuinely extend the scientific and intelligence frontier.
What’s missing is a neutral, open standard for evaluating these capabilities in real, production-scale intelligence development.
Today, we’re releasing OpenRSI-Index v0.1: evaluating whether AI can recursively improve itself and push the boundaries of intelligence and science, on production-scale clusters with 1k GPUs.
We turn fully open-source projects into autoresearch environments, with agent trajectories lasting 60+ hours. Building v0.1 took 100K+ H100-hours.
We’re building an ecosystem with and for the research community: let RSI benefit everyone, and let everyone shape RSI together.
We invite task contributors and compute partners to build this open benchmark with us - all contributors will be included as paper authors.
Shape RSI with us:
🌐 Website: https://t.co/unaB2yfQy4
🛠️ GitHub: https://t.co/vaRyltgbra
🤝 Contribute: https://t.co/PffRAhoIht
@OpenRSI It is my honor to contribute to this work in the RSI AutoResearch area. I look forward to making more contributions in the future, and I hope to see more contributions from the community as well!
1/ Blog
AI progress is limited not only by compute and data, but by the bandwidth of human insight.
The core ambition of RSI is to alleviate this bottleneck: by allowing agents to scale hypothesis generation and implementation, we can sweep a vastly larger method space than humans can explore alone.
Evaluating whether these loops can systematically surpass current methods is exactly what OpenRSI-Index is built for.
Research is never zero-sum: as agents handle the heavy lifting in the research loop, researchers are super-leveraged and get more room to chase crazier, paradigm-shifting ideas.
📖 Full write-up: https://t.co/VT8jmZITr5