Michael shares a great recap of the PyTorch conference and LLVM Developer meeting. He connects the two cultures and provides perspective on how both communities are touching similar problems from different directions. 👇
Finally had a chance to listen through this pod with Sutton, which was interesting and amusing.
As background, Sutton's "The Bitter Lesson" has become a bit of biblical text in frontier LLM circles. Researchers routinely talk about and ask whether this or that approach or idea is sufficiently "bitter lesson pilled" (meaning arranged so that it benefits from added computation for free) as a proxy for whether it's going to work or worth even pursuing. The underlying assumption being that LLMs are of course highly "bitter lesson pilled" indeed, just look at LLM scaling laws where if you put compute on the x-axis, number go up and to the right. So it's amusing to see that Sutton, the author of the post, is not so sure that LLMs are "bitter lesson pilled" at all. They are trained on giant datasets of fundamentally human data, which is both 1) human generated and 2) finite. What do you do when you run out? How do you prevent a human bias? So there you have it, bitter lesson pilled LLM researchers taken down by the author of the bitter lesson - rough!
In some sense, Dwarkesh (who represents the LLM researchers viewpoint in the pod) and Sutton are slightly speaking past each other because Sutton has a very different architecture in mind and LLMs break a lot of its principles. He calls himself a "classicist" and evokes the original concept of Alan Turing of building a "child machine" - a system capable of learning through experience by dynamically interacting with the world. There's no giant pretraining stage of imitating internet webpages. There's also no supervised finetuning, which he points out is absent in the animal kingdom (it's a subtle point but Sutton is right in the strong sense: animals may of course observe demonstrations, but their actions are not directly forced/"teleoperated" by other animals). Another important note he makes is that even if you just treat pretraining as an initialization of a prior before you finetune with reinforcement learning, Sutton sees the approach as tainted with human bias and fundamentally off course, a bit like when AlphaZero (which has never seen human games of Go) beats AlphaGo (which initializes from them). In Sutton's world view, all there is is an interaction with a world via reinforcement learning, where the reward functions are partially environment specific, but also intrinsically motivated, e.g. "fun", "curiosity", and related to the quality of the prediction in your world model. And the agent is always learning at test time by default, it's not trained once and then deployed thereafter. Overall, Sutton is a lot more interested in what we have common with the animal kingdom instead of what differentiates us. "If we understood a squirrel, we'd be almost done".
As for my take...
First, I should say that I think Sutton was a great guest for the pod and I like that the AI field maintains entropy of thought and that not everyone is exploiting the next local iteration LLMs. AI has gone through too many discrete transitions of the dominant approach to lose that. And I also think that his criticism of LLMs as not bitter lesson pilled is not inadequate. Frontier LLMs are now highly complex artifacts with a lot of humanness involved at all the stages - the foundation (the pretraining data) is all human text, the finetuning data is human and curated, the reinforcement learning environment mixture is tuned by human engineers. We do not in fact have an actual, single, clean, actually bitter lesson pilled, "turn the crank" algorithm that you could unleash upon the world and see it learn automatically from experience alone.
Does such an algorithm even exist? Finding it would of course be a huge AI breakthrough. Two "example proofs" are commonly offered to argue that such a thing is possible. The first example is the success of AlphaZero learning to play Go completely from scratch with no human supervision whatsoever. But the game of Go is clearly such a simple, closed, environment that it's difficult to see the analogous formulation in the messiness of reality. I love Go, but algorithmically and categorically, it is essentially a harder version of tic tac toe. The second example is that of animals, like squirrels. And here, personally, I am also quite hesitant whether it's appropriate because animals arise by a very different computational process and via different constraints than what we have practically available to us in the industry. Animal brains are nowhere near the blank slate they appear to be at birth. First, a lot of what is commonly attributed to "learning" is imo a lot more "maturation". And second, even that which clearly is "learning" and not maturation is a lot more "finetuning" on top of something clearly powerful and preexisting. Example. A baby zebra is born and within a few dozen minutes it can run around the savannah and follow its mother. This is a highly complex sensory-motor task and there is no way in my mind that this is achieved from scratch, tabula rasa. The brains of animals and the billions of parameters within have a powerful initialization encoded in the ATCGs of their DNA, trained via the "outer loop" optimization in the course of evolution. If the baby zebra spasmed its muscles around at random as a reinforcement learning policy would have you do at initialization, it wouldn't get very far at all. Similarly, our AIs now also have neural networks with billions of parameters. These parameters need their own rich, high information density supervision signal. We are not going to re-run evolution. But we do have mountains of internet documents. Yes it is basically supervised learning that is ~absent in the animal kingdom. But it is a way to practically gather enough soft constraints over billions of parameters, to try to get to a point where you're not starting from scratch. TLDR: Pretraining is our crappy evolution. It is one candidate solution to the cold start problem, to be followed later by finetuning on tasks that look more correct, e.g. within the reinforcement learning framework, as state of the art frontier LLM labs now do pervasively.
I still think it is worth to be inspired by animals. I think there are multiple powerful ideas that LLM agents are algorithmically missing that can still be adapted from animal intelligence. And I still think the bitter lesson is correct, but I see it more as something platonic to pursue, not necessarily to reach, in our real world and practically speaking. And I say both of these with double digit percent uncertainty and cheer the work of those who disagree, especially those a lot more ambitious bitter lesson wise.
So that brings us to where we are. Stated plainly, today's frontier LLM research is not about building animals. It is about summoning ghosts. You can think of ghosts as a fundamentally different kind of point in the space of possible intelligences. They are muddled by humanity. Thoroughly engineered by it. They are these imperfect replicas, a kind of statistical distillation of humanity's documents with some sprinkle on top. They are not platonically bitter lesson pilled, but they are perhaps "practically" bitter lesson pilled, at least compared to a lot of what came before. It seems possibly to me that over time, we can further finetune our ghosts more and more in the direction of animals; That it's not so much a fundamental incompatibility but a matter of initialization in the intelligence space. But it's also quite possible that they diverge even further and end up permanently different, un-animal-like, but still incredibly helpful and properly world-altering. It's possible that ghosts:animals :: planes:birds.
Anyway, in summary, overall and actionably, I think this pod is solid "real talk" from Sutton to the frontier LLM researchers, who might be gear shifted a little too much in the exploit mode. Probably we are still not sufficiently bitter lesson pilled and there is a very good chance of more powerful ideas and paradigms, other than exhaustive benchbuilding and benchmaxxing. And animals might be a good source of inspiration. Intrinsic motivation, fun, curiosity, empowerment, multi-agent self-play, culture. Use your imagination.
Recuerda que, ahora que ha vuelto @LaLiga, vuelven también los bloqueos a webs.
Si quieres entrar a @elOrdenMundial, te jodes.
Cualquier cosa, a @Tebasjavier.
Introducing Genie 3, the most advanced world simulator ever created, enabled by numerous research breakthroughs. 🤯
Featuring high fidelity visuals, 20-24 fps, prompting on the go, world memory, and more.
🔥 ¿Aire acondicionado de ricos y pobres?
El 75-85% de los hogares ricos lo tiene; pero solo el 40-55% de los pobres.
Una brecha que reduce tu productividad e impide que los niños aprendan. Y que mata. Aquí los datos 👇
Para que veais cómo se mueve ya esta gente, solo diré que el otro día nos llegó un burofax de @LaLiga a @elOrdenMundial.
En el mejor de los casos están disparando al bulto intentando meter miedo.
En el peor, están presionando a medios de comunicación.
Y al Gobierno se la suda.
Mis sitios web —y muchos miles más— están bloqueados en España cuando hay fútbol.
👏 ¡Un fuerte aplauso al Juzgado de lo Mercantil n.º 6 de Barcelona, a Telefónica y a todos los inútiles guionistas de este astracán, que tratando de mitigar un problema han creado otro mucho mayor!
Esto iba a pasar: los bloqueos de @LaLiga y @Telefonica están provocando pérdidas de ingreso en un montón de negocios online a los que les tumban la web cada vez que hay fútbol.
Aquí Tebas tiene carta blanca para hacer las barrabasadas que quiera.
https://t.co/VqNBT9QY80
Estimado @OdonElorza2011:
He leído con atención tu propuesta de «un twitter público gestionado por la UE». Y aunque entiendo tus argumentos y coincido contigo en que esto es —entre otras cosas— un cenagal de ruido y manipulación, discrepo radicalmente de tu planteamiento.
Estos son mis argumentos, que comparto muy respetuosamente en aras de un debate constructivo:
1️⃣ El problema de fondo no es Musk: es la internet de las plataformas. Internet es, por diseño, una red distribuida a la que cualquiera puede conectar nodos para servir o descargar contenido. Nadie debe pedir ni puede otorgar permiso, pues no hay una autoridad central. Nadie controla —ni, mucho menos, modera— lo que cada nodo publica. Los protocolos de comunicación entre los nodos son documentos públicos, elaborados por grupos de trabajo a los que cualquiera puede sumarse y cuyas deliberaciones son públicas¹.
Sobre dos de estos protocolos —el par HTML sobre HTTP— surge en los años 90, y luego eclosiona, la web. Que es una invención fundamentalmente europea, aunque esto importe poco en una red global como internet. Y desde los años 2010 vemos un desplazamiento de los contenidos desde la web hacia las plataformas.
Las plataformas son YouTube, Facebook, X, GitHub, TikTok, WhatsApp, Telegram, Reddit, WeChat, Discord, Twitch… Jardines que nada tienen que ver ni con la web ni con el diseño originario de internet: son espacios con una autoridad centralizada con facultades para conferir o retirar el acceso, cada una con su respectiva política de moderación de contenidos, sus generalmente opacos algoritmos de promoción de contenidos, sus propios y a menudo subrepticios intereses…
El problema fundamental de hoy no es Elon Musk: es el desplazamiento que en los últimos 15 años hemos hecho desde una web descentralizada por diseño a un oligopolio de plataformas con un poder concentrado en muy pocas manos. Y esto no se resuelve trasladando el control a un utópico «organismo independiente del gobierno y los poderes económicos» (sic) —tal cosa sería un mero intento de cambio del centro de gravedad—, sino justamente descentralizándolo de nuevo².
2️⃣ El problema de la moderación de contenidos no tiene solución. ¿Qué es una injuria? ¿Y un bulo? ¿Qué son exactamente el «tecnofascismo», el «trumpismo» o el «ciberpopulismo» que mencionas? ¿Quién decide cuáles son los «valores y principios» y la «ética democrática» que reclamas para una nueva plataforma pública de debate? Ni siquiera hay un consenso claro en torno a qué es la pornografía: un pecho desnudo en un retrato artístico es amonestado en Facebook pero, en cambio, tolerado en X.
No cabe aquí explicar por qué es imposible objetivar unos criterios universales de moderación de contenidos³. Basta recorrer la historia de puntillas para notar que los libros prohibidos un siglo son las obras más audaces del siguiente; que las más escandalosas ideas de un tiempo son la vanguardia del venidero; que las herejías que una época censuró son el faro de la generación posterior. La verdad y la mentira —como el bien y el mal— son un sistema relativista. Los marcos morales se desplazan en el tiempo⁴.
Y como no puede haber unos criterios objetivos y universales para discernir contenidos, tampoco puede haber ningún utópico «comité de sabios independientes», ningún algoritmo, ninguna inteligencia artificial que resuelva un problema que, sencillamente, no tiene solución.
3️⃣ El sector público no sabe, ni debe, hacer otra plataforma. Más rápida que la luz solo parece la velocidad con que el sector público olvida sus propios fiascos digitales. ¿Qué fue de Quaero, la iniciativa pública europea para «competir con Google»? Presentada por el tándem Chirac-Schröder con fuegos artificiales, Quaero recibió entre 2004 y 2010 cientos de millones de euros del contribuyente europeo para elaborar una tecnología paneuropea de búsqueda web⁵. Un esfuerzo público del que nada más se supo.
¿Qué fue del «Kelifinder» de la ministra Trujillo⁶? ¿Y de las distribuciones autonómicas de GNU/Linux que surtieron como setas de cada gobierno autonómico en 2005 y 2006, solo para dejarse discretamente morir de inanición poco después de apagados los focos de las ruedas de prensa? ¿Y de la app Radar Covid y los cuatro millones de euros que costó? ¿Y de la web de Renfe, ese desagüe? ¿Y de LexNet, el víacrucis digital de un sistema de justicia donde los expedientes se desplazan en carritos de supermercado? ¿Y de los innumerables marketplaces públicos⁷? ¿Qué fue del cacareado «infojobs público»? ¿Y del «idealista público»?
La retahíla de fracasos de la digitalización con el dinero de todos es tan dilatada que sonroja pensar ahora en un «twitter público». ¿Por qué iba a funcionar un «twitter de la UE» cuando ni siquiera compañías especialistas como Meta o Google, con enormes recursos financieros, han conseguido articular una masa crítica mínimamente viable en torno a sus Threads o Google+? Un espíritu prudente debe rechazar tal pretensión, aunque solo sea por el principio de eficiencia y economía que debe regir el gasto público (art. 31.2 de la Constitución Española).
Más que a la utopía, el mundo público debería mirarse urgentemente al espejo: su fantasioso enfoque a la privacidad en línea, por ejemplo, ha infestado la web de unos «avisos de cookies» que técnicos y usuarios destestamos por igual, y que se han demostrado del todo ineficaces salvo en una cosa: su capacidad de molestar⁹.
👉 Mucho más podría argumentar, aportando datos, hechos y experiencias, más allá de opiniones que solo se sostienen en un romanticismo idealista. Pero no cabe aquí. Como resumen, solo diré que «un twitter de la UE» sería un intento imposible de poner a quien no sabe a resolver el problema equivocado.
(Siguen, en el mensaje de abajo, las referencias).