one nn learns every atari game at once in realtime from scratch in one hour on a 4090. 56 seconds of gpu time per game. no pause or reset or memory peeking just 60fps color images. i was an rl skeptic two weeks ago. now i don't know. just trained this today.
"Nanda drastically underestimates, Nanda drastically underestimates the progress of recent circuit sparsity work, Jacobian SAEs, singular learning theory, and SPD/APD-style sparsification. Leo Gao. Yeah I read that too.”
Is there a "notricks halting problem" which excludes self referential programs? Like the actual halting problem. I might've asked this question before ten years ago but i can't remember the answer