We are starting Intent Lab, building an autonomous team we call "fleet" that turns intent into production software. Today we are sharing some early results: the fastest GLM5.2 inference engine, one shot database creation, and a fully verified agent filesystem.
https://t.co/CtwtypE0rH
@ anyone
DM me if u need help with (literally) anything.
I try to pay forward the help many have given to me.
Just be genuine haha and I’ll help where I can.
Episode 1 of @OriginsInc with Jack Altman (@jaltma):
Jack is the founder of Lattice, a multi-billion dollar company. Now, he manages Alt Capital, a $150M early-stage venture fund.
He’s one of the most thoughtful and genuine people we’ve had the privilege of sitting down with, and a good friend of ours and @OriginsInc.
We discuss Jack’s early upbringing, time in college, and going from Princeton to investment banking. Soon quitting, he invested as a venture capitalist with @sama and joined the fastest-growing YC startup... which later failed.
This led to the formation of Lattice, a multi-billion dollar company, and now Alt Capital.
Timestamps:
0:00 Intro
1:38 Childhood and growing up
5:23 Parents, school and brothers
13:33 Getting into banking
15:29 Jack's wife
17:40 Career & investing
28:52 Building Lattice
35:25 Strengths, weaknesses, and tradeoffs
36:55 Hobbies and children
41:48 Alt capital
43:42 What drives you?
45:50 Today’s youth
49:33 Intuition
50:35 Predicting the future
52:25 Cold email, rejection, and competing
58:27 Kindness and helpfulness
01:00:15 Day trading and stocks
01:02:00 What people don’t know about Jack
GPTFast 0.3.1 is here🚀!
-Stabilized GPTQ
-Customized W4A16 Kernels for MatMul which are 𝟯𝟬% 𝗳𝗮𝘀𝘁𝗲𝗿 than nn.Linear on RTX 3050
-In practice, GPTQ doesn't work well with spec decoding
-Benchmarks needed on more GPUs to figure out perf bottlenecks https://t.co/See9UhnPDk
- Have you thought about contributing to vllm, but found it scary? Have no fear, vllmini is here! 🚀
- Me and @andre_slav03 built vllm from the kernels on the bottom to FastAPI on top
- Educational for now, hope to make it prod ready soon
- GitHub link: https://t.co/MgJMdxTXPz
Performance does seem to be saturating however - GPTQ INT4 quantization speeds up inference speeds by 9x but INT8 already increases it by 8.5x. This would seem to indicate that performance is memory bound and not compute bound. Future versions should address this issue. 2/n
@joao_gante@maximelabonne Very interesting, I will definitely take a look at the implementation 👀
I also wanted to ask about your thoughts on this: https://t.co/chzgphGl50. If you are familiar with the architecture, you can use meta-programming to modify any model directly 😼