Huge thanks to the devs behind https://t.co/vtsnwII4Hv for making Triton work on Windows π
Thanks to their work, we could use torch.compile with a Triton backend β even though official Triton doesnβt support Windows yet.
Open-source wins again! π§π₯
Been digging into tokenization while following @karpathy Zero-to-Hero path.
Wrote my own BPE tokenizer from scratch β added regex-based grouping, chunked encoding (like GPT-2), and learned a ton along the way.
If you're curious:
π https://t.co/MPLZiLNT1t
I have 3 tickets to the MLOps World conference
Either in person in Toronto or online
Want to win one?
πΈοΈ Follow me
πΈοΈ Retweet this tweet
π https://t.co/XyOikVMmWR
Winners selected randomly and announced on Wednesday
Good luck π
Learn MLOps in a free online course from @DataTalksClub!
- Processes
- Training models
- Serving models
- Monitoring
- Best practices
More info here: https://t.co/Y68jLzbI9N
@karpathy Great plan! You're probably already overwhelmed with invitations. But in case you are open to meeting and making new friends, it would be super cool to meet in Germany for beer or tea :)
@yoavgo One more addition. Here https://t.co/wUrlaiiade is their network bandwidth explanations. I think, one should be able to achieve a better performance than 20 minutes for 200 GB. One more point for trying support :)
@yoavgo Even though they have quite extensive docs, it is not that obvious from the first glance to me what would be the bandwidth between gcp and vm. You may try to reach they support, they might have some better options to share.
@yoavgo I think, if the data is somehow transformed on the regular basis, it would be better to save it on persistent (mountable to VM) disc and then do all the transformations.