That’s the irony of it - her books are philosophy made palatable by stories, which many people confuse for the reverse.
She also somehow made society’s unrelatable characters feel human, while showing what its supposed saints become when you follow their ideas to their second-order effects.
The first experimental evidence of recursive self-improvement (RSI).
Autoresearching the autoresearch agent for eight days.
The result beats the harness we hand-tuned for two years, on held-out benchmarks: 🧵(1/7)
Your autoresearch needs its own Weights & Biases.
We’ve turned Weco into an observability tool that lets you monitor, analyze, and share autoresearch runs. Here's what it can do: 🧵(1/4)
We're excited to announce that @BingchenZhao, who built the predecessor of AutoResearch, has joined @WecoAI full-time!
Bingchen is the first author of LLMSpeedrunner at Meta FAIR, which ran the automated research loop on @karpathy's NanoGPT, which later evolved into NanoChat and the speedrun community where AutoResearch operates today.
Weco has been committed to ML research automation for 2.5 years, starting with AIDE. We're super pumped by how large an impact AIDE has had, topping @OpenAI's MLE-Bench and @METR_Evals' RE-Bench, and becoming a foundation for AI Scientist v2, AIRA-Dojo, and LLMSpeedrunner itself.
And AutoResearch, with AIDE's simple greedy discard/keep loop reaching a mass audience, is really building consensus that the empirical research loop can and should be automated. We're excited to keep pushing this frontier, not just as a concept but seriously bringing it to the real world, and materially accelerating the knowledge generation of humanity.
In case you want to run AutoResearch this weekend:
It costs ~$300 for 85 experiments using Claude Code (opus).
A quick guide to autoresearch ~60 experiments for free:
1. Use the mac/local GPU fork:https://t.co/wRnQgdsomi
2. Use weco to get some free credits: `pipx install weco` → `weco setup claude-code` Or simply give this doc to your Claude Code agent: https://t.co/aEebJguABo
- You’ll get $20 in free credits
3. Tell your coding agent to run weco optimization for val_bpb on https://t.co/vlD5lDTbnI.
4. Tell your coding agent to use gemini-3-flash-preview, you should get about 60 free experiments.
- For better performance, use gemini-3.1-pro-preview (~15 free experiments).
5. You can watch the progress on this nice dashboard: https://t.co/WJ2UawSCfL
@dexhunt3r@tobi Yeah, @WecoAI's been doing did this for a couple of years now. Glad the industry is finally catching up to back then while Weco's pushing the frontier: https://t.co/DhCLkqKGKf
Spending precious time fine-tuning your model?
Weco is the alternative to fine-tuning. It iterates through thousands of prompts to improve performance on any model, on any task.
Built for teams pushing on production metrics or benchmarks.
Thrilled to announce Weco has raised an $8M seed led by @GoldenVentures to build self-evolving software!
Our technology has already been used by frontier labs like OpenAI, Meta, Google and Sakana AI.
We’re making every codebase a living experiment that learns to beat itself:
Why is everyone suddenly talking about Super Intelligence, even though AGI isn’t here yet?
Here's my take: “super‑human on some axes” will arrive sooner than “human‑level on all axes”:
Since GPT-3, the view in frontier AI research assumed a linear progression: Specialized Weak AI → AGI → ASI, where intelligence is a single scalar (like IQ).
But reality is painting a different picture.
As @ylecun put it around 2 years ago: "Research in AI is like climbing a mountain range in the fog. Every time we reach what we thought was the summit, the clouds lift a bit and we realise we were only on a ridge with higher peaks ahead."
With more data points having come in the last a few years it's getting clear: intelligence isn't one-dimensional. Some aspects of human cognition are proving far harder to replicate than others.
What LLMs have mastered:
- Extraordinarily broad knowledge bases
- Processing speed that dwarfs human capability (often underappreciated)
- Language mastery, e.g understanding, generation, multilingual fluency
- Code generation (if you think it's also a type of language)
- Well-formatted reasoning (math, coding challenges)
Rapid progress in:
Multi-step decision making, tool use (agentic capabilities)
The stubborn challenges:
- Memory: Humans encode experiences in neural weights; LLMs rely on context windows, more like note-taking than true memory
- Vision: Two years after GPT-4V, SOTA multimodal models still struggle to read analog watches without tools
- Creativity: This remains contentious. I lean conservative here as it might emerge after we solve other problems like agentic capabilities + memory.
- Embodied Intelligence: Physical world interaction remains elusive
So what does the future hold?
My prediction: Certain cognitive capabilities of large generative models will rapidly surpass human levels (some arguably already have). With proper scaffolding, we'll see superhuman performance in domains with stationary environments and clear rewards, think of super human algorithm discovery already done by AlphaEvolve.
These systems will feel like magic upon release. But as with all technology, we'll quickly normalize their capabilities and treat them as tools.
We may approach AGI capable of "replacing all white-collar jobs," but the timeline is likely longer than optimists suggest.
The path to intelligence isn't a single mountain: it's a complex landscape with peaks we're only beginning to glimpse.
Hey, @cursor_ai! Adding a selectable Pin icon when hovering over a context (selected or about to be) would help a lot - especially for temporarily including things like style guides, so new prompts use them by default
@elevenlabs Would be cool for @Apple to replace their Dictation tech with something of this level. It would make a huge difference for @cursor_ai development
I used to spend weeks in trial-and-error loops building deep learning models, until we built AIDE to handle that work for us. Now I can tackle more than 20 ML problems at once and train 1,000+ models in parallel. It’s incredibly empowering! See how we’re rethinking machine learning engineering in our latest arXiv paper.
The entire internet era has been about transmitting information. AI will take computer science to the next level, by generating new knowledge through trial and error.
LLMs following the scientific methodology can already automate R&D.
Check the paper of AIDE!