To take ASI seriously is to accept a weak-to-strong premise: human intelligence can build a process that eventually produces intelligence beyond itself. RSI is the last piece of that puzzle of ASI. It serves as the key to scaling on insights: human insights are limited by bandwidth, while agents could scale the generation of ideas and sweep a far larger region of method-space.
As the early project initiator, we are fully aware that today, everyone is excited about RSI, yet the barriers to participation also keep everyone outside the door. We believe that a model's RSI capabilities intrinsically come from researchers, independent labs, and domain teams alike, and we think the benefits should ultimately return to everyone who builds models.
This is why we take our name as OpenRSI-Index. We aim to keep RSI open through shared platforms and tools, so more people can participate and benefit. We could shape RSI together, set RSI’s standards, and challenge frontier models with our own related real research work.
For this goal, we build RSI-Anything, a Human-AI collaboration pipeline. Through about one hour of conversation, the pipeline could help turn a real research question into a runnable autoresearch environment. We want researchers and agents to solve these problems together, sharing new insights, methods, and results with the community.
Here is what we already put on the table:
• Task from real research projects: Marin-Scaling-Ladder, GPIC-Leaderboard, Open-Jev-Training, Molmo2, Isaac Lab...
• Ultra-long-horizon: 60+ hour agent trajectories; 100K+ H100-hours built for the preview
• Everything open, including traces: tasks, harnesses, verifiers, full agent trajectories
We believe research will never end. A game turns zero-sum only when the pie is too small to share. However, research is definitely the field with the highest ceiling there is. What shifts is the mindset: it frees researchers to find and formulate the crazier, more valuable, more exciting problems in the world.
To take ASI seriously is to accept a weak-to-strong premise: human intelligence can build a process that eventually produces intelligence beyond itself. RSI is the last piece of that puzzle of ASI. It serves as the key to scaling on insights: human insights are limited by bandwidth, while agents could scale the generation of ideas and sweep a far larger region of method-space.
As the early project initiator, we are fully aware that today, everyone is excited about RSI, yet the barriers to participation also keep everyone outside the door. We believe that a model's RSI capabilities intrinsically come from researchers, independent labs, and domain teams alike, and we think the benefits should ultimately return to everyone who builds models.
This is why we take our name as OpenRSI-Index. We aim to keep RSI open through shared platforms and tools, so more people can participate and benefit. We could shape RSI together, set RSI’s standards, and challenge frontier models with our own related real research work.
For this goal, we build RSI-Anything, a Human-AI collaboration pipeline. Through about one hour of conversation, the pipeline could help turn a real research question into a runnable autoresearch environment. We want researchers and agents to solve these problems together, sharing new insights, methods, and results with the community.
Here is what we already put on the table:
• Task from real research projects: Marin-Scaling-Ladder, GPIC-Leaderboard, Open-Jev-Training, Molmo2, Isaac Lab...
• Ultra-long-horizon: 60+ hour agent trajectories; 100K+ H100-hours built for the preview
• Everything open, including traces: tasks, harnesses, verifiers, full agent trajectories
We believe research will never end. A game turns zero-sum only when the pie is too small to share. However, research is definitely the field with the highest ceiling there is. What shifts is the mindset: it frees researchers to find and formulate the crazier, more valuable, more exciting problems in the world.
OpenRSI just launched OpenRSI-Index v0.1 today. Here's what you need to know.
OpenRSI-Index is a new open benchmark for measuring recursive self-improvement, or RSI, in frontier AI model development. The idea is to test whether AI agents can push past human-designed methods to genuinely extend science and intelligence on their own, at real production scale rather than in toy setups.
The benchmark turns fully open-source projects into autoresearch environments. Agents get let loose on these environments in trajectories that run 60-plus hours, operating on clusters with 1,000 GPUs, closer to how real frontier labs actually train and research.
Building v0.1 already consumed over 100,000 H100-hours. The project also has a contribution pipeline called RSI-Anything, which walks outside researchers through turning their own research question into a packaged task in about an hour, covering baseline, evaluation, and budget.
OpenRSI says it is being built as an open ecosystem, inviting task contributors and compute partners, with all contributors credited as paper authors on the eventual write-up.
Key numbers:
- v0.1 preview released today
- Agent trajectories run 60+ hours
- Tested on clusters with 1,000 GPUs
- 100,000+ H100-hours spent building v0.1
- RSI-Anything task pipeline takes about 1 hour
Code and docs are live now on GitHub and at https://t.co/OA2o5Z5RhM.
Let RSI benefit everyone and let everyone shape RSI.
Excited to be one of the project initiators helping the design of the signiture task and organization. I also think most of the research can be formulated as 1) hill climbing a current task, and 2) formalize a new task. We would like to share the process with everyone, rather than locking them only in some places.
Must-read paper from Google on self-improving agent harnesses.
If you auto-optimize your agent's harness, your eval score can go up while the agent gets worse on real tasks.
This paper shows how to prevent that.
Of five harness-evolution methods compared on agentic workspace tasks, RRSI scored the lowest on the tasks it evolved against and highest on all three out-of-distribution benchmarks.
Automated harness evolution proposes edits to prompts, control flow, tools and memory, keeps the ones that raise the score, and repeats.
The authors show this overfits the training tasks. Meta-Harness reached 93.0 on the Harvey LAB evolve split but gained only 0.3 to 1.5 points on JobBench, GDPval and APEX-Agents.
RRSI adds regularization on both sides of the loop.
The proposer gets an edit budget that shrinks over time and is pushed toward directions it has not tried. A critic rejects benchmark-specific edits, and a pruner removes edits that are too small, too costly or no longer useful.
RRSI scored 90.5 on the evolve split and gained 3.5 to 4.7 points on the three held-out benchmarks. In the ablation, unregularized evolution used 3.80M tokens per trial against 2.42M for RRSI. With Gemini 3.5 Flash, RRSI raised Terminal-Bench 2.1 from 64.6 to 78.7 and carried a 2.2-point gain over to SWE-bench Verified.
Paper: https://t.co/SlfjDg96VI
Chat with Paper: https://t.co/gBotiH6Jfq
Open recursive self-improvement is the next frontier of AI to genuinely expand scientific boundaries.
To scale OpenRSI-Index v0.1, we are calling for Domain Leads and Compute Partners!🚀
Join us in shaping the future of RSI! 🌐
As AI systems enter a recursive self-improvement loop, the central question is whether it can systematically move beyond human-designed methods to genuinely extend the scientific and intelligence frontier.
What’s missing is a neutral, open standard for evaluating these capabilities in real, production-scale intelligence development.
Today, we’re releasing OpenRSI-Index v0.1: evaluating whether AI can recursively improve itself and push the boundaries of intelligence and science, on production-scale clusters with 1k GPUs.
We turn fully open-source projects into autoresearch environments, with agent trajectories lasting 60+ hours. Building v0.1 took 100K+ H100-hours.
We’re building an ecosystem with and for the research community: let RSI benefit everyone, and let everyone shape RSI together.
We invite task contributors and compute partners to build this open benchmark with us - all contributors will be included as paper authors.
Shape RSI with us:
🌐 Website: https://t.co/unaB2yfQy4
🛠️ GitHub: https://t.co/vaRyltgbra
🤝 Contribute: https://t.co/PffRAhoIht
Today, we’re launching ResearcherScout to help you hire AI researchers at conferences.
Application companies are becoming the new neolabs. Now everyone wants to build a great research team.
But you can’t judge a researcher by their LinkedIn profile. You need to understand their papers, their actual work, and meet them in person.
One user described the tool as a "lifesaver."
ECCV main conference starts from Sep 10th, try it now: https://t.co/GxtCIDPnfO.
OpenRouter: 40 months → $7B+ reported acquisition.
It took them ~2.5 years to reach $10M ARR.
OrcaRouter: 2.5 months → on track for $10M ARR in-month. 🐋
A few more numbers from our first 10 weeks ↓
Introducing dots3-note preview — a small but mighty step toward long-horizon agency in real life.
🔹 280B MoE with 16B active parameters, a 512K context window, and multimodal understanding across text, vision, and audio
🔹 Introduces TEMPO, a new RL approach for long-horizon agent training through self-critiquing and test-time-scaled value estimation
🔹 Built to reason, explore unfamiliar environments, update memory over time, and combine multimodal perception with coding and tool use to solve complex tasks
🔹 Open weights on Hugging Face, alongside two open benchmarks for real-life agents: VibeSearchBench and VibeLifeBench
Competitive with much larger models across reasoning, agentic, and multimodal evaluations.
🔗 Tech blog: https://t.co/kMJtfvB1s7
🔗 Model weights: https://t.co/Lxwbgjc8wk
🔗 Github: https://t.co/QAAeYBTULP
Introducing FrontierPhysics: benchmark for e2e frontier physics research
We built SkillsBench and got 180+ citations in 6mo. People have been asking next project to contribute to
Led by @bingran_bry - Physics PhD @UCBerkeley, Nature author, advised by Prof. Haeffner and more 🧵
🚀🚀🚀 Give AI agents personas!
New research on world simulation from Harvard, MIT, and Stanford. We introduce 𝗠𝗮𝘁𝗿𝗔𝗜𝘅, a population-scale simulated-user evaluation infra for testing AI systems and digital products.
𝗠𝗮𝘁𝗿𝗔𝗜𝘅: 𝗦𝗶𝗺𝘂𝗹𝗮𝘁𝗶𝗻𝗴 𝘁𝗵𝗲 𝗪𝗼𝗿𝗹𝗱 𝘄𝗶𝘁𝗵 𝟴.𝟯 𝗕𝗶𝗹𝗹𝗶𝗼𝗻 𝗣𝗲𝗿𝘀𝗼𝗻𝗮 𝗔𝗴𝗲𝗻𝘁𝘀 🌍
🔗 Project and Community: https://t.co/trvAPK5H0J
📄 Paper: https://t.co/MktjtmriJJ
🎬 Demo: https://t.co/hadWrE9QlH
💻 GitHub Repo: https://t.co/50AUdIHX3O
👥 Persona 8B: We created 8.3B persona records at the scale of the global population, across 1,290 dimensions spanning background, psychology, capabilities, behavior, and lifestyle.
🧪 MatrAIx Playground: Supports persona-agent evaluations across four environment types: Survey, AI Chatbot, Web, and App.
📚 MatrAIx Applications: Offers 1,000+ evaluation tasks across 25+ domains, including Commerce, Software, Finance, and Healthcare.
🤝 Join our open-source research community (https://t.co/dzcLWYwqO2) of 200+ contributors, including 40+ from OpenAI, Anthropic and Google DeepMind!
@XiaominLi98@YuexingHao
#AI #LLM #Agents #Persona #Evaluation #Infra #UserSimulation #Harvard #MIT #Stanford #OpenAI #Anthropic #GoogleDeepMind #xAI
Coding benchmarks are saturating. AI4Research is the next frontier.
Thrilled to see our MLS-Bench (https://t.co/yExIXjAjkv) becoming the first AI4Research benchmark to gain broad recognition and adoption across the community.
Congratulations to the team!
CORAL is heading to COLM 2026! 🪸🎉
Thrilled to share that our work, “CORAL: Towards Autonomous Multi-Agent Evolution,” has been accepted to COLM 2026!
What happens when AI agents move beyond rigid workflows and start collaborating, organizing, accumulating knowledge, and evolving together? Check out CORAL and drop us a ⭐ if you find it interesting: https://t.co/WjUJlG88VX
See you at COLM 2026—let’s grow the reef together! 🪸
📖Paper: https://t.co/TH3BGG4cWU
#agentic #llms #selfevolvingagent #multiagent #autoresearch #alphaevolve #colm
🎉 Congrats to OpenAI on GPT-5.6's strong performance on Agents' Last Exam!
We created Agents' Last Exam (ALE) to provide realistic, reproducible evaluations and continuous public measurements of frontier AI agents on economically valuable, real-world work across broad domains. ALE currently includes 1,500+ expert-sourced, long-horizon tasks spanning 55 non-physical occupations, contributed by 300+ experts from more than 100 institutions. All tasks are grounded in real professional work and evaluated with fully verifiable, outcome-based grading.
It's exciting to see ALE becoming an important guidepost for frontier agent development—helping the community track progress, better understand emerging capabilities, and build AI systems that are increasingly reliable, capable, and useful in the real world.
A huge thank you to the hundreds of experts and collaborators who have contributed to ALE. This is truly a community effort. We're continuing to expand the benchmark with new tasks and domains, and we'd love for more experts to contribute. Together, we can build an even stronger foundation for measuring progress toward capable, trustworthy Agentic AI.
Submit your contribution: https://t.co/S07s1hjG4L
Excited to see our JobBench was adopted in the Muse Spark 1.1 release as a benchmark for measuring professional agentic workflows. JobBench is targeting to enhance humans rather than replace them with GDP values like GDPVal/ Remote Labour Index. See https://t.co/9M56H5za9o!
GPT-5.6 is here.
Sol is an incredible model, and Terra/Luna provide great performance at lower price.
Great at coding, knowledge work, cybersecurity, and science with fewer tokens and at lower cost.
https://t.co/XXRz1HmMsv
On Agents' Last Exam, GPT‑5.6 Sol sets a new high of 53.6, eclipsing Claude Fable 5 (adaptive) by 13.1 points.
At medium reasoning, it beats Fable 5 by 11.4 points at roughly one-quarter the estimated cost. GPT‑5.6 Terra and Luna also outperforms Fable 5 at around one-sixteenth the cost.