Can an agent explore a new environment, learn its causal structure, and keep improving without updating its model weights?
We introduce RSIAgent, a framework for recursive self-improvement through autonomous exploration. Using Kimi-K3 and GLM-5.3 as base models, RSIAgent outperforms GPT-6 Astra on both OSWorld 2.0 and Agents’ Last Exam.
RSIAgent decides what to explore, executes tasks, verifies outcomes, and consolidates stable action-condition-outcome relationships into memory for future use.
With the underlying model weights fixed, RSIAgent achieves:
- 78.98% Partial Score on OSWorld 2.0 (0808 offline), compared with 72.60% for GPT-6 Astra
- 84.82% on Agents’ Last Exam (Near-term), compared with 82.26% for GPT-6 Astra
We call this Scaling Experience. Agents can continue improving by acquiring, verifying, and reusing their own experience while the underlying model weights remain fixed.
Links in the reply below.
For AI to work with us, it needs to understand us
Today, we're introducing Persimmon, the first large-scale model designed to realistically simulate how people talk and interact
Today, we're sharing new research on Solaris, our first Interface World Model.
Solaris is a new kind of operating system that generates interactive interfaces frame by frame, in real time, with no code. We find that Solaris outperforms frontier LLMs when generating new interfaces, across structural similarity and information retention. Read more and request early access at the link below.
AI power users are showing up outside tech. The fastest-growing Codex adopters since February:
Legal: 108x
Sales: 41x
Recruiting: 41x
Marketing: 26x
Healthcare: 24x
Charts of the Week: https://t.co/7YT2BXvuUu
Depth-aware light injection in TypeGPU
I got a 448x448 monocular depth model down to ~8 ms on my M4 Pro across ~250 dispatches, which is fast enough to use in realtime :D
Since the inference is written directly in TypeGPU, I can just feed the depth buffer straight into the lighting pass. It never has to leave the GPU or go through any extra synchronization/interop step
Inference, lighting and draw all go through the same command encoder.
Stripe for AI ending up with Stripe feels almost too perfect!
So many models, so much infra needed to make them play nicely together. Good year to be the plumbing.
Stripe has finalized an agreement to acquire OpenRouter, a startup that helps companies switch between artificial intelligence models, for more than $7 billion, according to people familiar with the matter. https://t.co/NZRt1tYhX5
Introducing Cues: Voice & Gesture control for Matic
We raised $115M and spent 9 years to make the world’s first intuitive home robot
Say you spilled coffee: Point to the spill and say, "Hey Matic, clean this" and it will hear, see, locate in 3D, and go clean on its own.
Matic comes with a lot of features:
1. Say "Hey Matic" - it locates your voice, turns, and looks at you
2. Say "Hey Matic, follow me" and start walking. Matic will follow behind
3. Say "Hey Matic, go clean the living room". Since it knows your house map, it navigates and just does it
It's so easy, a 5 year old and an 80 year old can use it and it understands 75 different languages.
Matic has 8x the airflow, specialised cleaning algorithms for rugs, corners, toekicks, mopping, etc and cleans better than any other robot vacuum.
Also keeps improving with software updates.
13,000 families use and love Matic.
WIRED magazine gave it a 10/10 (the only hardware to receive this rating in a decade)
Buy yours at https://t.co/nLfY5pchxB and if you don't love it after 6 months, we'll give you a full refund.
To celebrate our launch, we're cleaning 300 homes with Matic in San Francisco and New York City.
Comment "Matic" below, we'll send you the link to sign up and come to your doorstep to clean your home with Matic.