I don't have too too much to add on top of this earlier post on V3 and I think it applies to R1 too (which is the more recent, thinking equivalent).
I will say that Deep Learning has a legendary ravenous appetite for compute, like no other algorithm that has ever been developed in AI. You may not always be utilizing it fully but I would never bet against compute as the upper bound for achievable intelligence in the long run. Not just for an individual final training run, but also for the entire innovation / experimentation engine that silently underlies all the algorithmic innovations.
Data has historically been seen as a separate category from compute, but even data is downstream of compute to a large extent - you can spend compute to create data. Tons of it. You've heard this called synthetic data generation, but less obviously, there is a very deep connection (equivalence even) between "synthetic data generation" and "reinforcement learning". In the trial-and-error learning process in RL, the "trial" is model generating (synthetic) data, which it then learns from based on the "error" (/reward). Conversely, when you generate synthetic data and then rank or filter it in any way, your filter is straight up equivalent to a 0-1 advantage function - congrats you're doing crappy RL.
Last thought. Not sure if this is obvious. There are two major types of learning, in both children and in deep learning. There is 1) imitation learning (watch and repeat, i.e. pretraining, supervised finetuning), and 2) trial-and-error learning (reinforcement learning). My favorite simple example is AlphaGo - 1) is learning by imitating expert players, 2) is reinforcement learning to win the game. Almost every single shocking result of deep learning, and the source of all *magic* is always 2. 2 is significantly significantly more powerful. 2 is what surprises you. 2 is when the paddle learns to hit the ball behind the blocks in Breakout. 2 is when AlphaGo beats even Lee Sedol. And 2 is the "aha moment" when the DeepSeek (or o1 etc.) discovers that it works well to re-evaluate your assumptions, backtrack, try something else, etc. It's the solving strategies you see this model use in its chain of thought. It's how it goes back and forth thinking to itself. These thoughts are *emergent* (!!!) and this is actually seriously incredible, impressive and new (as in publicly available and documented etc.). The model could never learn this with 1 (by imitation), because the cognition of the model and the cognition of the human labeler is different. The human would never know to correctly annotate these kinds of solving strategies and what they should even look like. They have to be discovered during reinforcement learning as empirically and statistically useful towards a final outcome.
(Last last thought/reference this time for real is that RL is powerful but RLHF is not. RLHF is not RL. I have a separate rant on that in an earlier tweet
https://t.co/RMIpFPVpuM)
4/ As people turn to Google for deeper insights and understanding, AI can help us get to the heart of what they're looking for. We're starting with AI-powered features in Search that distill complex info into easy-to-digest formats so you can see the big picture then explore more
1/ In 2021, we shared next-gen language + conversation capabilities powered by our Language Model for Dialogue Applications (LaMDA). Coming soon: Bard, a new experimental conversational #GoogleAI service powered by LaMDA.
https://t.co/cYo6iYdmQ1
#AlphaCode—a new #AI system for developing computer code developed by @DeepMind—can achieve average human-level performance in solving programming contests.
Learn more this week in Science: https://t.co/baygVUiHmi
Applications for the first-ever Google PhD Fellowships for students in Latin America open today, along with applications to support early-career professors through Research Scholar. Read more about our investments in the Latin American research ecosystem ↓https://t.co/JSnaAqkukA
If you are out of ideas, go for a walk.
This paper found walking (whether outdoors or a treadmill) increased key types of creative thinking for over 80% of undergraduates. The reasons are not fully clear, but there seem to be direct effects on the brain. https://t.co/0EDfpzyf8j
Two important breakthroughs from @GoogleAI this week - Imagen Video, a new text-conditioned video diffusion model that generates 1280x768 24fps HD video. And Phenaki, a model which generates long coherent videos for a sequence of text prompts. https://t.co/nTs67r21Sf
Since 1969 Strassen’s algorithm has famously stood as the fastest way to multiply 2 matrices - but with #AlphaTensor we’ve found a new algorithm that’s faster, with potential to improve efficiency by 10-20% across trillions of calculations per day! https://t.co/nLvFbEDBuO
“The equivalent of a James Webb Space Telescope for biology."
#AlphaFold will accelerate scientific research and discovery in ways we're only beginning to appreciate.
Excited to announce that we've set up the first Google Research Hub in Australia! We've got a great leadership team, with Grace Chung, Peter Bartlett, and @stevemblackburn, and we're hiring! More at https://t.co/cpyA22AzY1. @GoogleAI
How can we rigorously evaluate the anticipated risks of language models (LMs)?
Our team identifies six “characteristics” of LM-generated text which can help guide the design of more rigorous benchmarks. Learn more: https://t.co/2mLweMCxK4 1/
Today we present the first demonstration of a provable exponential advantage in #QuantumMachineLearning over classical algorithms that is robust even on today's noisy hardware. Learn all about it at https://t.co/BmAlLqWK8E