Splitting an 8B model across eight H100s reduces weight-streaming time from 4.9 ms to 0.6 ms per token. Measured decode latency improved by less than 2×. That distance was the subject of this talk. Thanks @BengaluruSys
Here's what went down at July's Edition!
1. @iamsaradindu covered the three principal strategies for partitioning a model across GPUs and the communication cost associated with each
(1/2)
So performance becomes a question of how well those systems are used, not how fast the code is. I took ideas from "Performance Hints"(https://t.co/CaahDCK3Sb) and translated the underlying principles into concrete Python ML patterns: https://t.co/x86PgcExWI
"Python is slow" is one of the most repeated and least useful things said about ML engineering. In a well-built system, Python shouldn't be doing the work at all — NumPy runs C, PyTorch runs CUDA, tokenizers run Rust.
Will build a working context engine: agents that read, write, and reason over shared structured context, making memory and decision history a first-class part of the system at #MCPDevSummit
https://t.co/o0eotsa89G
Knowing how to write assembly is a skill you should learn, and these guys have a great resource for you!
I've debugged 10,000 lines of assembly for every line I've ever written... but writing assembly from scratch is a core computer science skill, I believe: even if you never use it.
Will you use it to write new code? Maybe not. But when you get dropped into a call stack without source code, at least you won't have to ask your Dad for help!
And the next thing you should do after learning to write ASM is to get a crisp understanding of what your C code actually compiles to. Like a switch() statement is often a jump table, but if you've never debugged one...
Fantastic work. Is there a way to support this through voluntary donations? This should be done for every state & MP and be part of future ballots.
Excited to speak at this year's Pydata Global on ML unlearning. I started working on ML unlearning after the first competition launch at 2023 NeurIPS. Looking forward to it.
https://t.co/SGHahsJOP1
‼️ AND SO IT BEGINS! PYDATA GLOBAL IS OFFICIALLY UNDERWAY ‼️
Get ready for three days of data brilliance, innovation, & worldwide collaboration. The wait is finally over; check your email for registration details! 🌐
Let the data festivities COMMENCE! 🎉https://t.co/YcprNjKUnp
Really excited to be speaking at @gdgcblr Community Day. I will be speaking on a topic in ML that I am interested in, interpretability in ML production systems. #GoogleCloud#GCCDBLR2023
📷
🎉 Join us at #GCCDBLR2023 for an incredible lineup of speakers across tracks! Get ready to be inspired by industry experts and gain valuable insights. Don't miss this engaging day of sessions and networking. See you there! 🌟
#GoogleCloud#GCCDBLR2023
About 2 years ago, when a private equity firm bought StackOverflow for $1.8B, I was unsure if this would be a good exit. 'Surely StackOverflow should be worth more?' I thought.
However, knowing what we know now: it looks like the perfectly timed exit.
https://t.co/q4dNmV817z
Awesome presentations by Madhusudan, Sathish and Arpana. Reconnecting with enthusiastic developer community on a Sunday after 2 years of #remotework and #wfh
At Google Cloud Community Day
#GCCDBLR@gdgcblr#googlecloudnext
Data quality metrics:
Summary statistics:
1 mean
2 median
3 variance
4 skewness
5 min max range
6 percentage of null, zero and uniques
7 standard deviations
#GCCDBLR@gdgcblr@iamsaradindu
Managing li-ion batteries during and after their use is critical from sustainability pov, especially in a country like India where everything runs at scale and building that tech stack is interesting @nunamenergy https://t.co/6aeo4kPwOD
2nd life Audi e-tron batteries are to power electric rickshaws by #Nunam in India to enable women’s economic participation. Find out more about this project of Audi Environmental Foundation.
Get a test drive on @greentech_fest >> https://t.co/qpA8sHPuES
#Audi#sustainability