Karpathy's Autoresearch is bottlenecked by a single GPU. We removed the bottleneck.
We gave the agent access to our K8s cluster with H100s and H200s and let it provision its own GPUs. Over 8 hours:
• ~910 experiments instead of ~96 sequentially
• Discovered that scaling model width mattered more than all hparam tuning
• Taught itself to exploit heterogenous hardware: use H200s for validation, screen ideas on H100s
Full setup and results: https://t.co/03WnG9zmzO
@karpathy
7/ Beyond Human Data for LLMs - an approach for self-training with feedback that substantially reduces dependence on human-generated data; the model-generated data combined with a reward function improves the performance of LLMs on problem-solving tasks.
https://t.co/pnsP90EHJx
10/ Quip - compresses trained model weights into a lower precision format; combines lattice codebooks with incoherence processing to create 2 bit quantized models; significantly closes the gap between 2 bit quantized LLMs and unquantized 16 bit models.
https://t.co/w1u6JKHkrg
Today, we share our teams’ latest contributions, Phi-2 and promptbase.
Phi-2 outperforms other existing small language models, yet it’s small enough to run on a laptop or mobile device. https://t.co/wLhUeRsByL