Overdue announcement, I am a PyTorch Foundation Member Representative 🔥
Glad to be contributing to the community and hosting a PyTorch meetup at our London office on 16 September 🍻
We’re nearly at capacity, register soon to secure your spot👇
https://t.co/xtd4ZK7Vf8
Join us at Graphcore London for the London PyTorch Meetup!
📅 16 Sept, 6–9pm
🎤 Talks from Arsalan Uddin, Andrew Fitzgibbon and the wider PyTorch community
🍕 Food, drinks and networking
Register: https://t.co/K50lQnS9UZ
#PyTorch#LondonTech
11 billion parameters. On a phone.
We wanted to see whether Llama-3.2-11B-Vision-Instruct — far too large for a phone in its original form — could fit into a ~4 GB budget.
The answer: yes. Here's how 🧵
Spring is here and so is Papers of the Month! In this March edition, we cover Transformers without Normalisation, Compute Optimal Scaling of Skills, Overtrained Language Models Are Harder to Fine-Tune, and Multi-Domain Distribution Learning for De Novo Drug Design! 🧵
Our November Papers of the Month is now live.
This edition covers 4 papers, on "super weights", context-parallelism, scaling laws for precision and critical batch sizes. We provide our summaries and analysis of each. (🧵1/n)
https://t.co/CkQySUiNpO
At @raais this year, @sylvainviguier of @graphcore on the need to drive significant efficiency in AI, as training and energy costs spiral.
We’re back next year, 13 June 2025, join us!
raais(dot)co
Introducing `tandv` - a library for tracking and visualising the internal stats of your model. We hope this will help with low-precision, debugging and more.
(link in 🧵)
We've written a roundup of ICML and the papers we found interesting. For all those keen on sparsity, speculative sampling and schnitzel...
https://t.co/kqnDv1zS6Q
LLM model size competition is intensifying… backwards!
My bet is that we'll see models that "think" very well and reliably that are very very small. There is most likely a setting even of GPT-2 parameters for which most people will consider GPT-2 "smart". The reason current models are so large is because we're still being very wasteful during training - we're asking them to memorize the internet and, remarkably, they do and can e.g. recite SHA hashes of common numbers, or recall really esoteric facts. (Actually LLMs are really good at memorization, qualitatively a lot better than humans, sometimes needing just a single update to remember a lot of detail for a long time). But imagine if you were going to be tested, closed book, on reciting arbitrary passages of the internet given the first few words. This is the standard (pre)training objective for models today. The reason doing better is hard is because demonstrations of thinking are "entangled" with knowledge, in the training data.
Therefore, the models have to first get larger before they can get smaller, because we need their (automated) help to refactor and mold the training data into ideal, synthetic formats.
It's a staircase of improvement - of one model helping to generate the training data for next, until we're left with "perfect training set". When you train GPT-2 on it, it will be a really strong / smart model by today's standards. Maybe the MMLU will be a bit lower because it won't remember all of its chemistry perfectly. Maybe it needs to look something up once in a while to make sure.
After nearly a year of waiting, it was great to attend the @raais summit of 2024! The highlight for me was @ThoreG amazing AI+biology presentation on an exploration of resilience in the body and what aging is.
Photo: RAAIS
Once again @PunchdrunkInt did not disappoint. Loved their new immersive piece made of a narrated story in a confined space, lights that beg to be followed and powerful sonic atmospheres.