in case anyone is interested in how this was solved.
in the first big kernel, “xmma_gemm” tells us its a matmul and it seems like its using bf16 with fp32 accumulate
and the “sm90” tells us the architecture of the GPU being used. sm90 = H100 (sm80 = A100, sm70 = V100)
Fun puzzle of the day: given the PyTorch trace, identify model + parallelism strategy. Then, label every kernel in the image
I've found there's so much performance alpha in being able to walk through every kernel when tuning models
$1 prize to first solve
“targeting frontier LLM development […] pretraining pipelines, distributed training infrastructure, or ML accelerator design”
post-training snubbed. definitely someone on the pretraining team wrote this
When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model’s capabilities through methods such as prompt modification, steering vectors, and PEFT.
Anthropic estimated that this would affect approximately 0.03% of traffic.
“Hey Claude can you implement this?”
*Razzmatazzing*
“Absolutely! I created a Jira ticket for it in the backlog and will prioritize it in 2 weeks during the next sprint planning. Let me know if you’d like me to T-Shirt size that now or during the next backlog grooming session.”
AI labs are paying hundreds of thousands of dollars to buy email, Slack and Jira threads from dead startups as feedstock for ‘reinforcement learning gyms,’ which specialize in using defunct company data to build simulated work environments https://t.co/QuCk36wAPY
@dphuang2 cool project btw. any plans to share what the model predicts during the other majors this season? would be fun to follow alongside the prediction markets live
@dphuang2 > After R3, his lead had collapsed. Tied with Cameron Young. Kalshi: 36%. Model: 35%.
for fun, what was your own p(Rory win) after R3? I was pulling for him but thought for sure he would’ve needed better than -1 on sunday