Introducing The Darwin Gödel Machine: AI that improves itself by rewriting its own code
https://t.co/wEEB4LGPr0
The Darwin Gödel Machine (DGM) is a self-improving agent that can modify its own code. Inspired by evolution, we maintain an expanding lineage of agent variants, allowing for open-ended exploration of the vast design space of such “self-improving” agents.
Modern agentic systems, while powerful, remain static—once deployed, their intelligence remains fixed. We believe continuous self-improvement is key to the development of stronger AI capabilities. Our Darwin Gödel Machine is built from the ground up to enable AI systems that can learn and evolve their own capabilities over time, just as humans do.
On SWE-bench, DGM automatically improved its performance from 20.0% to 50.0%. Similarly, on Polyglot, the DGM increased its success rate from an initial 14.2% to 30.7%, significantly outperforming representative hand-designed agents.
Learn more about our approach in our technical report: https://t.co/kDNWFgCI6C
This work was done in collaboration with Jeff Clune (@jeffclune)’s lab at UBC, and led by his PhD students Jenny Zhang (@jennyzhangzt) and Shengran Hu (@shengranhu), together with Cong Lu (@cong_ml) and Robert Lange (@RobertTLange).
Code: https://t.co/RcYLd22TB5
New tutorial! I spent 3 weeks realizing flow-matching/rectified flows can be viewed in a simple way that end-runs the usual pages of math:
"Basic physics provides a 'straight, fast' way to get up to speed with flow-based generative models"
Colab included! https://t.co/NHscQI7qCX
Farewell to Achim Szepanski (1957-2024), a member of the influential German experimental P16.D4 in 1982, but perhaps best known for founding the legendary German techno label Force Inc. Music Works, with the sub labels Mille Plateaux and Ritornell.
🧪🤖 Super excited to share Synthemol: a #genAI that not only designs new molecule drugs but tells us how to build each molecule in lab.
It designed 6 novel antibiotics that we synthesized in Ukraine and experimentally validated https://t.co/owgIIxvfV0 @NatMachIntell 🧵
Minimal implementation of the Tree of Attacks (TAP) LLM jailbreaking from @robusthq:
https://t.co/NXJcFKJLGF
- Cleaned + updated the prompts
- Uses OpenAI, Mistral, and TogetherAI APIs
- Refactored the leaf branching strategy
Meta presents Self-Rewarding Language Models
paper page: https://t.co/ZAh4ZotyCL
Fine-tuning Llama 2 70B on three iterations of our approach yields a model that outperforms many existing systems on the AlpacaEval 2.0 leaderboard, including Claude 2, Gemini Pro, and GPT-4 0613
Paste this post into ChatGPT 4. 😊
Fast-forward ⏩ alignment research from @GoogleDeepMind ! Our latest results enhance alignment outcomes in Large Language Models (LLMs). Presenting NashLLM!
I know everyone is excited about Mixtral and the new Hyena models - but @ContextualAI just dropped a pile of cool new models and a new alignment framework
https://t.co/Mmn3Vvff3T
🧵 (1/n)
👉 Introducing QuIP#, a new SOTA LLM quantization method that uses incoherence processing from QuIP & lattices to achieve 2 bit LLMs with near-fp16 performance! Now you can run LLaMA 2 70B on a 24G GPU w/out offloading!
💻 https://t.co/eM8Br2CqL3
Project #2: LLM Visualization
So I created a web-page to visualize a small LLM, of the sort that's behind ChatGPT. Rendered in 3D, it shows all the steps to run a single token inference. (link in bio)
🎉 Exciting update: 📚🌐 the "wikimedia/wikipedia" dataset is now available with the latest @Wikipedia dump across all 323 languages in Parquet format: https://t.co/S7lYPbcADh A valuable resource for language models. Let the training begin! 🤖 #AI#NLP@Wikimedia@WikiResearch