Tom Tellez on Stride Development-
During lunch in Houston years ago Tom shared a packet of notes he had, and much of it was not included in his recent book.
What was fascinating was his thoughts on acceleration and top speed, much different than the contemporary work today. 🧵
Tom Tellez on Stride Development-
During lunch in Houston years ago Tom shared a packet of notes he had, and much of it was not included in his recent book.
What was fascinating was his thoughts on acceleration and top speed, much different than the contemporary work today. 🧵
These things just pluck numbers from the aether.
My dog spent the whole night whining and scratching at the door. I feel like I haven't slept at all, was woken at least once an hour: 91/100
.@IsraelOlatunde5 has smashed his own Irish 100m record by clocking 10.12 (1.7m/s) to win at the NEB Open in London.
Previous Irish record was the 10.17 he ran in the 2022 European final.
📷 @sportsfile
Adeleke asked if there's any positive to take from finishing 4th in an Olympic final at just 21.
"No, that’s not positive. Some people come here to participate and are just happy to be at the Olympics and to call themselves an Olympian. I knew what I was capable of."
@sportsfile
10.71 winning his heat, 10.65 (1.0) winning his semi 🔥. Both performances bettering Aaron Sexton’s @AthleticsNI U18 record from 2017. Final tomorrow at 19.50
In 2019, OpenAI announced GPT-2 with this post:
https://t.co/jjP8IXmu8D
Today (~5 years later) you can train your own for ~$672, running on one 8XH100 GPU node for 24 hours. Our latest llm.c post gives the walkthrough in some detail:
https://t.co/XjLWE2P0Hp
Incredibly, the costs have come down dramatically over the last 5 years due to improvements in compute hardware (H100 GPUs), software (CUDA, cuBLAS, cuDNN, FlashAttention) and data quality (e.g. the FineWeb-Edu dataset). For this exercise, the algorithm was kept fixed and follows the GPT-2/3 papers.
Because llm.c is a direct implementation of GPT training in C/CUDA, the requirements are minimal - there is no need for conda environments, Python interpreters, pip installs, etc. You spin up a cloud GPU node (e.g. on Lambda), optionally install NVIDIA cuDNN, NCCL/MPI, download the .bin data shards, compile and run, and you're stepping in minutes. You then wait 24 hours and enjoy samples about English-speaking Unicorns in the Andes.
For me, this is a very nice checkpoint to get to because the entire llm.c project started with me thinking about reproducing GPT-2 for an educational video, getting stuck with some PyTorch things, then rage quitting to just write the whole thing from scratch in C/CUDA. That set me on a longer journey than I anticipated, but it was quite fun, I learned more CUDA, I made friends along the way, and llm.c is really nice now. It's ~5,000 lines of code, it compiles and steps very fast so there is very little waiting around, it has constant memory footprint, it trains in mixed precision, distributed across multi-node with NNCL, it is bitwise deterministic, and hovers around ~50% MFU. So it's quite cute.
llm.c couldn't have gotten here without a great group of devs who assembled from the internet, and helped get things to this point, especially ademeure, ngc92, @gordic_aleksa, and rosslwheeler. And thank you to @LambdaAPI for the GPU cycles support.
There's still a lot of work left to do. I'm still not 100% happy with the current runs - the evals should be better, the training should be more stable especially at larger model sizes for longer runs. There's a lot of interesting new directions too: fp8 (imminent!), inference, finetuning, multimodal (VQVAE etc.), more modern architectures (Llama/Gemma). The goal of llm.c remains to have a simple, minimal, clean training stack for a full-featured LLM agent, in direct C/CUDA, and companion educational materials to bring many people up to speed in this awesome field.
Eye candy: my much longer 400B token GPT-2 run (up from 33B tokens), which went great until 330B (reaching 61% HellaSwag, way above GPT-2 and GPT-3 of this size) and then exploded shortly after this plot, which I am looking into now :)
Another year of successful Final Year Project presentations with over 250 attendees! Our students showcased a brilliant range of innovative work.
We thank our students, staff & partners (@RAEngNews @EngineersnBiz, Women in STEM) for their engagement and support.
#Engineering
The Northern Ireland Biomedical Engineering Society Annual Symposium will be held at the Ulster University Belfast Campus on Thursday the 23rd of May 2024. Registrations Open Now!
#NIBEC2024@NIBES_official
Teaching is an incredible joy and also a great chance to learn.
Full article, magazine front page, and intranet homepage (not sure deserved, but a true honor)! Thank you @MayoClinic for highlighting this work. What a great place to work! 😃
📖 Magazine: https://t.co/xYuwf1vPT6
World Athletics are planning to trial a new format to the long jump 🏟️
This is after data from the World Championships in Budapest highlighted a third of attempts were no-jumps 🇭🇺
Instead of a take-off board, there will be a take-off zone and it will be trialled at lower-level competitions this season 🌍
Jumps would then be measured from the front of the take-off foot within that zone 🌬️
Speaking on @AnythingbutF, World Athletics CEO Jon Ridgeon stated: "We’ll measure from where the athlete takes off to where they land in the pit.
“That means every single jump counts. It adds to the jeopardy and drama in the competition.
"We’ll spend this year testing it in real life circumstances with very good athletes. If it doesn’t pass testing, we’ll never introduce it.
"At the same time we’re working out ways we can get instant results so you don’t have to wait 20-30 seconds before the result pops up."