it's a spectrum between 'ai systems can literally help us in no way' to 'ai systems literally develop their hardware/software entirely', and we're inching closer to the latter
I feel like the term "recursive self-improvement" has grown from a truly dangerous thing -- an AI system that is sufficiently smart and well-equipped that it can autonomously improve *itself* -- to "any feedback loop where any AI system is useful for building future AI systems"?
LLM agents have demonstrated promise in their ability to automate computer tasks, but face challenges with multi-step reasoning and planning. Towards addressing this, we propose an inference-time tree search algorithm for LLM agents to explicitly perform exploration and multi-step planning in interactive web environments.
It is the first tree search algorithm for LLM agents that shows effectiveness on realistic and complex web environments: on the challenging VisualWebArena benchmark, applying our search algorithm on top of a GPT-4o agent yields a 39.7% relative increase in success rate compared to the same baseline without search, setting a state-of-the-art success rate of 26.4%. On WebArena, search also yields a 28.0% relative improvement over a baseline agent, setting a competitive success rate of 19.2%.
I like the new Sonnet. I'm frequently asking it to explain ML papers to me. Doesn't always get everything right, but probably better than my skim reading, and way faster.
Automated alignment research is getting closer...
With @Harvard, we built a ‘virtual rodent’ powered by AI to help us better understand how the brain controls movement. 🧠
With deep RL, it learned to operate a biomechanically accurate rat model - allowing us to compare real & virtual neural activity. → https://t.co/GaToq3AWTQ
Google DeepMind CEO Demis Hassabis on AI accelerationists:
“They don’t actually fully understand the enormity of what’s coming. Because if they did — I’m very optimistic we can get this right — but only if we do it carefully and take the time needed to do it.”
Aravind Srinivas says the role of CEO is highly replaceable and future AI like GPT-5 or 6 could perform CEO tasks better by continuously processing information and making decisions.
Model interpretability techniques allow us to understand and control what models do. Here's how Sparse Autoencoders give us general fine grained control on image models:
bro dalle 3 getting wild wtf. why is it so good? these dont even look ai generated they look like theyre from a movie or something. dalle 3 always getting iteratively better, openai just never says anything about it.
Virtually all large models today contain huge matrices, and these dominate their compute. By incorporating structure in these matrices, we can improve the performance/compute tradeoff!