Demain, à 16 heures sur @franceculture, nous ne ménagerons pas nos méninges, en compagnie de Richard Lévy.
Un cerveau, comment ça marche ? via @franceculture https://t.co/ph00EMkdZL
«A 2020 Vision of Linear Algebra» six brief videos for teaching and learning linear algebra by… yasss! Him! A 85 years old William Gilbert Strang! ❤️❤️❤️
I hope to have just a fraction of your stamina, dear Gilbert!
I’ll try to keep opening eyes!
https://t.co/wXMJfchynv
Je viens de numériser un livre d'Elias et Scotson assez dur à trouver si jamais ça vous intéresse. C'est une étude d'un quartier anglais qui essaye de comprendre sociologiquement comment se créent les logiques de groupe et d'exclusion.
Je vous mets un lien juste en dessous.
A new tokenizer is introduced for LLMs:
https://t.co/3P0Z2blvBN
Idea: Instead of merging tokens by frequency (BPE), optimize the tokenizer directly for maximizing average token length, yielding longer, more efficient tokens.
Results: 14–18% fewer tokens, faster training & inference, and better downstream accuracy across (small) GPT-style models.
“So much of physics comes down to understanding geometry. And often in surprising ways.” —Jonathan Sorce, a theoretical physicist at Princeton University https://t.co/953Kc4Pduj
Deep learning is not a collection of black-box tricks, contrary to what many believe. It can be learned as a principled engineering discipline.
This latest edition of Deep Learning with Python is my best attempt so far at teaching it. It focuses on building deep intuition -- the theory and mental models -- alongside all of the practical programming patterns.
You will learn this on a truly modern stack: Keras 3 as a high-productivity, framework-agnostic API, and preferably JAX for SotA performance and scalability (PyTorch and TensorFlow are also fully supported).
Une lecture en soutien à Henry Laurens pour analyser les enjeux de l'annulation du colloque sur la Palestine & l'Europe au College de France : "Le Passé imposé" montre la nécessité de l'histoire pour dépasser l'affrontement des mémoires.
Vertigineux et plus que jamais nécessaire.
Why and how does gradient/matrix orthogonalization work in Muon for training #LLMs?
We introduce an isotropic curvature model to explain it. Take-aways:
1. Orthogonalization is a good idea, "on the right track".
2. But it might not be optimal. [1/n]