Vibe hardware is here.
I gave Astra my credit card and asked for a Teenage Engineering-style mini DJ controller.
It generated a concept image, sourced parts, read Chinese datasheets, built a CAD model, ordered everything, then made a Blender animation showing how to assemble it.
We're releasing the arXiv paper on AIDE².
The RSI system where AI research agents improve their own research efficiency.
It includes new results on transfer across models and comparisons with more AI research agents: [1/4]
The arXiv paper on AIDE² is out.
AIDE² improved its own research agent beyond the version we hand-tuned for two years, and the gains hold on benchmarks the outer loop never saw.
New in the paper are transfer across models and comparisons with more AI research agents.
Our autoresearch agent improves the foundational layer of LLMs: its pretraining data.
And we are seeing RSI transforming every layer of the AI stack.
Great work @vivian_yanmy and the Weco team!
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
One of our autoresearch runs sat flat for 60 steps. One message got it moving again.
Everyone is pushing research agents toward full autonomy, but what helps most is being able to step in when a run goes wrong, without breaking the loop. (1/6)
I had a lot of Fable tokens to use up before my weekly reset, so I made this live 3D map of London with Three.js
Every train, bus, boat and plane is real and live right now!
- Tube, bus and riverboat data from TfL
- National Rail trains from Darwin live departure boards
- ADS-B for planes and helicopters
- AIS feed for boats and ships
- Map data from Overture and OpenStreetMap
Trains and buses have no GPS feed, so their positions are inferred from arrival countdowns and departure boards, then animated along the track/route geometry
OpenAI ran a hiring challenge, but the top candidate was one they couldn’t hire: our autonomous research agent, Aiden.
In Parameter Golf, Aiden ran for 22 days, and out-outperformed all 1,016 other researchers: 🧵 (1/8)
Is autoresearch really better than classic hyperparameter tuning?
We did experiments comparing Optuna & autoresearch.
Autoresearch converges faster, is more cost-efficient, and even generalizes better: 🧵(1/6)
Autoresearch has been out for 2 weeks. The community is trying to apply it to everything with a measurable metric, here are some successful attempts: 🧵 (1/6)
Your autoresearch needs its own Weights & Biases.
We’ve turned Weco into an observability tool that lets you monitor, analyze, and share autoresearch runs. Here's what it can do: 🧵(1/4)