Timelapse - 0.5 (3 hrs)
Didn’t plan a long session. Just poked at Jev and wrote a tiny BPE.
- Read what a System One model is and tried jev in the playground.
- Implemented a small Python BPE to see how a basic tokenizer actually works
Timelapse-00 (8 hrs)
- Continuing my project on B200 matmul
- Went through how clusters and 2CTA MMA work
- Decided on two improvement paths: 2CTA and a staged epilogue
- Went through Qwen3.5 model architecture; planning to write an inference engine for it.
Building a DeepSeek-V4-Pro style inference lab from scratch on NVIDIA B200 with CUDA/PTX.
v0 is live: single-token sliding-window attention core.
Goal: own the full inference dataflow - decode, KV cache, gather, kernels, architecture details (SWA/CSA/HCA, mHC, MoE).
Repo: https://t.co/kcKZGG6gy9
Finally finished a v1 BF16 GEMM on B200 (raw CUDA/PTX)
~75% of cuBLAS on 2048 / 4096 / 8192.
Still more learning and implementing than beating anything.
A lot of the structure is from Paul Chan's writeup: https://t.co/KxtApqmi2l
Repo: https://t.co/W71JNxtqws
Been a while - time to get hacking!
Back on my raw CUDA DeepSeek-style inference project. Got a small vertical slice of the sliding-window attention decode core running.
Nothing fancy yet.
Next: validate the target shape, measure on B200, then optimize.
>o)
(_>
@0xSero Humanity advanced the fastest when we openly shared knowledge. Now AI lets everyone not just compress existing knowledge but also create new knowledge - potentially the key to building a Dyson sphere. And besides, a game is way more fun with 100 pro players than just 2 pros.
Btw, this video lecture on H100 by @_PrateekShukla_ is really amazing (link below) - it walks you through PTX, CUTLASS, and core GPU logic + fundamentals in detail.
Highly recommended even if you’re working on Blackwell or any other architecture.
\(^o^)/
https://t.co/msMjLBQAhi
One really awesome thing about open-weight models: they’re so cheap that I have zero mental block about burning tokens or making my prompts super efficient.
This is actually pushing me to experiment more and integrate AI deeper into my workflow.
Pretty fun ngl :D