I just wrote an analytical report on Linear Attention and Sparse Attention, thoroughly discussing their respective strengths and weaknesses from both mathematical and engineering perspectives.
https://t.co/lAJlSSFXrM
Yesterday, NVIDIA published a blog post about the Rubin architecture, and I conducted some detailed analysis based on the preview version of the PTX 9.4 Instruction Set guide.
https://t.co/Ap7KPoV7nV
BREAKING: Alibaba says its new AI model is "second only" to Anthropic’s Fable 5.
The Chinese tech giant says its Qwen3.8 Max has 2.4T parameters, comparable to leading frontier AI models and second only to Fable 5.
The ranking is Alibaba’s own claim and has not yet been independently verified.
There's a homework assignment in the Anthropic interview about architectural optimization with VLIW+SIMD. It's a really interesting problem.
I spent a few hours on it and have already gotten it down to 987 cycles.
https://t.co/Zt86eNXheb
It's a super complex task. My plan^3 is so convoluted that even I'm getting dizzy trying to wrap my head around it. There's just no way I can describe it to the model in plain language. https://t.co/vqXzIOK2rM
Recently, I used an AI agent to develop a library for scatter host-to-device memory copying—without writing a single line of code myself, The final implementation achieved several times the performance of cudaMemcpy.https://t.co/KYgdEFUeH7