Thrilled to be partnering with @ValsAI on their RSI Index: a benchmark for whether models can do the work of an AI researcher.
How close are we to true RSI?
Results below, plus the full experiment ledgers in marimo notebooks so you can watch each model try to build its own successor ⬇️
AI safety is back in the spotlight, driven by concerns about recursive self improvement. We built the first third party benchmark to measure just how close AI is building its own successor.
In collaboration with @marimo_io and @CoreWeave we built the The RSI Index, which runs every frontier model through the same AI research tasks and scores each one against the strongest published results. So far, the models can do the work, but is far from the human frontier.
For our cheminformatics competition, you can now win an @Apple Mac Mini M6. You have 1 month left to take a dataset and build a marimo notebook that brings cheminformatics to life. Competition details below.
My Socratic learning skill hit 50 GitHub stars ⭐️ milestone!
https://t.co/ulhZ4waPOg
Why need a whole skill when simply promoting “socratic” would do a job? My skill keeps a learning journal for you and gauge your understanding when you come back to it again.