Last week we welcomed the PyTorch London community to our London office for an evening of shared ideas and thoughtful discussion.
Thanks to Arsalan Uddin, Andrew Fitzgibbon and Akshat Tripathi for sharing their insights - and to everyone who brought the conversation to life.
11 billion parameters. On a phone.
We wanted to see whether Llama-3.2-11B-Vision-Instruct — far too large for a phone in its original form — could fit into a ~4 GB budget.
The answer: yes. Here's how 🧵
Our latest work uses theory from the '50s to figure out how to design weight quantisation formats for LLM inference.
It's called Optimal Formats for Weight Quantisation and has just hit arXiv.
1/6
Spring is here and so is Papers of the Month! In this March edition, we cover Transformers without Normalisation, Compute Optimal Scaling of Skills, Overtrained Language Models Are Harder to Fine-Tune, and Multi-Domain Distribution Learning for De Novo Drug Design! 🧵
@EzProgramming@allenholub I think the better way of thinking about it is in the reference frame of the club head: the ball is approaching it at the speed of the club, and bounces off it in the opposite direction.
For a perfectly elastic collision it would go at 2x the club speed (in our reference frame)
@OpenAI Finally! A benchmark that seems to capture the vibe of the last 9 months of "clause sonnet 3.5 is best for coding".
Looking forward to testing other models against it!
Summary: https://t.co/djqFBtyRr3
Finally, the Phi-4 paper presents a rather different FLOPs angle: spending compute in the data-generation process to create higher quality data, leading to “student” models that (in some domains) out-perform their “teachers”.
Each month our team writes up summaries and analysis of our favourite ML papers. For December we cover:
The Byte Latent Transformer, Large Concept Models, Memory Layers & Phi-4 — all grouped under the title "Spend Your FLOPs Wisely". Here's what we made of them 🧵
Our November Papers of the Month is now live.
This edition covers 4 papers, on "super weights", context-parallelism, scaling laws for precision and critical batch sizes. We provide our summaries and analysis of each. (🧵1/n)
https://t.co/CkQySUiNpO
We’ve also updated our license to allow developers to use the outputs from Llama models — including 405B — to improve other models for the first time.
We’re excited about how this will enable new advancements in the field through synthetic data generation and model distillation workflows, capabilities that have never been achieved at this scale in open source.
New blog post by @douglasahorr breaking down the components of a modern transformer LLM.
Specifically written for programmers and casual ML enthusiasts, to give an understanding & minimal implementation of what's going on under-the-hood.
https://t.co/COoQiJAYs8