Announcing one year of LLM inference metadata traces, with 6.12 billion requests.
We hope this dataset can support research on real-world LLM serving workload understanding, system design and infrastructure optimization. Explore the dataset and learn more: https://t.co/5nFSp2PPaa
Driven by our great graduate student William Nixon and
in collab with @jon_durbin@airesearch12@chutes_ai
OpenAI researcher "Alisa Liu" had 57 interviews before joining OpenAI, and then She open-sourced her entire study notes job/learning process.
- her LLM study notes
- her math interview notes
- a full honest write-up of the job process
If you are preparing for research scientist / MTS roles, this is the highest-signal material available right now.
You can use her notes or her topic list to study on your own. That’s rare. Don’t waste it.
we launched the most comprehensive ai performance engineering repo in the world
now we'll be posting every single resource
this is Wafer's ai performance engineering series
save this to keep up with the series. links in thread 🧵
part 3: "Intro to CUDA C++" from NVIDIA's CUDA Programming Guide.
NVIDIA covers the execution model, memory movement, and correctness checks behind CUDA programs:
- kernel launches, grid dimensions, and the organization of threads into blocks.
- thread indexing and work assignment, including bounds checks for inputs that aren't multiples of the block size.
- unified memory and explicit memory management, including control over data placement and transfers between CPU and GPU.
- asynchronous kernel execution and synchronization before the CPU uses GPU results.
- shared memory and block-level synchronization for threads that need to exchange data and coordinate their work.
- runtime initialization and the setup costs that can affect measurements of the first runtime calls.
- error handling for kernel launches and execution, including failures that surface in later API calls.
- checking GPU results against a CPU implementation with a floating-point tolerance.
the worked examples connect these concepts in a complete vector addition program, showing how to divide the work, manage its memory, and check the results before moving on to more complex kernels.
figure from An Even Easier Introduction to CUDA (Updated)
we launched the most comprehensive ai performance engineering repo in the world
now we'll be posting every single resource
follow and save to keep up with the series. links in thread 🧵
part 4: Programming Massively Parallel Processors: A Hands-on Approach
the authors explain how GPU hardware executes parallel programs and how memory access and work distribution affect performance.
for an AI performance engineer, the book connects those fundamentals to the algorithms and optimization techniques used to diagnose slow kernels and decide what to change.
it covers:
- CUDA programming and GPU execution, including CPU and GPU cooperation, multidimensional grids, warps, scheduling, and synchronization. the vector-addition, image-processing, and matrix-multiplication examples show how to assign work to threads and reason about execution efficiency.
- memory hierarchy, coalescing, tiled matrix multiplication, thread coarsening, and occupancy. the authors explain how to reduce memory traffic while accounting for register and shared-memory usage, then provide a checklist for identifying a computation's bottleneck.
- convolution and stencils, including constant memory, caching, shared-memory tiling, and register tiling. these examples show how neighboring outputs can reuse input values, reducing repeated memory accesses when computing over arrays and grids.
- histograms, reductions, and prefix sums, including atomic operations, privatization, and work efficiency. the authors show how to reduce contention, limit divergence, and combine partial results without adding unnecessary memory traffic or computation.
- merging and sorting, including input partitioning, tiled merge, radix sort, and merge sort. these chapters connect algorithm choice and thread-to-data mapping to memory coalescing and the distribution of work across the GPU.
- sparse matrices and graph traversal, including sparse matrix-vector multiplication and breadth-first search. the authors compare storage formats and parallelization strategies to show how memory access, control divergence, and contention affect performance.
- deep learning, including perceptron inference and backpropagation, convolutional neural networks, a CUDA convolutional-layer inference kernel, convolution expressed as matrix multiplication, and cuDNN. this chapter connects the book's GPU programming techniques to the implementation of neural-network computations.
- case studies in MRI reconstruction and electrostatic potential mapping. the authors work through parallelism, loop transformations, memory layout, and validation, including how scatter and gather approaches change the cost of a computation.
- computational thinking and parallel algorithm design, including algorithm selection and problem decomposition. these provide a method for finding parallel work and choosing an implementation around the computation's requirements.
- CUDA streams and heterogeneous clusters, including MPI communication and CUDA-aware MPI. a distributed stencil example shows how to overlap communication with computation and coordinate work across GPUs.
- dynamic parallelism and advanced CUDA practices, including GPU-launched kernels, zero-copy memory, unified memory, and profiling and debugging tools. the book examines how kernels launch work and access data, including the limitations that can affect execution efficiency.
- numerical considerations, including floating-point representation, rounding, arithmetic accuracy, and numerical stability. the appendix explains how arithmetic and algorithm choices affect the reliability of computed results.
the book develops these ideas through worked kernels and applications, then connects them to neural-network computation in its deep-learning chapter.
for an AI performance engineer, that makes the material useful for reasoning about kernel execution, reducing memory traffic, and checking numerical results.
this Stanford paper is f*cking insane
they just compressed the entire hedge fund playbook into a 17-page pdf.
Stanford put out the complete Hidden Markov Model framework that quants at firms like Jane Street and Two Sigma are known to run, and released it for free.
the crazy part is this isn't some watered-down summary, it's the actual mechanics behind models these desks keep locked up internally.
most people assume this stuff never leaves institutional walls.
this one hands you the framework directly, no gatekeeping.
bookmark it and read before someone takes it down.
Woah… a new report on Anthropic vs. White House negotiations over Fable says things got so heated that Trump apparently wanted them JAILED. (per Politico)
> government tests Fable
> Treasury approves its release
> 2 days after launch Amazon does a jailbreak
> White house tells Anthropic to take it down
> Dario says jailbreaks happen for all models and it’s not catastrophic
> tries to explain why
> Bessent: “Well, I don’t think that’s really necessary, Dario”
> Dario said he couldn’t just shut down Fable on an hour’s notice
> a frustrated Trump reportedly says “I want to send them to jail”
> White House blocks access for all foreign nationals, including Anthropic’s OWN employees
> Dario asks if that means they have to take the model down
> Lutnick: “Yes. That’s the point.”
> Anthropic takes it offline worldwide
> they negotiate
> Fable returns after 19 days
wild… china might have another serious frontier lab on its hands, it was released 2 days ago..
> 600B MoE, around 30B active
> 1M context
> benchmarks put it around Kimi K3 / Grok 4.6 territory
> it’s much cheaper than most frontier models
> it’s biggest interest is agentic coding + tool use
> multimodal , text + image input.
i still wanna see if this benchmaxxed or not but on paper crazy how a lab founded in 2023 is already getting this close to the frontier..
JUST IN: Q* has been solved.
Welcome to the frontier of RL scaling.
KL-Regularized Policy Optimization for Critic-Free Agentic Reinforcement Learning
https://t.co/OWTl0J42xV
From GPO, SPPO, and RPG to BPO and Score Centering, we have finally arrived at the grand finale of RL science and Q*.
the way Óscar Restrepo in Un Poeta (2025) goes on a vitriolic rant about how José Asunción Silva is the best ever Colombian writer, far better than Garcia Marquez 'who was thirsty for recognition'—is lowkey how i saw Jean-Joseph Rabearivelo juxtaposed with the negritude poets
Set in a dystopian Los Angeles on the eve of the new millennium, Kathryn Bigelow’s Strange Days (1995) is one of the rare gems of the ’90s. The film imagines a world where SQUID technology lets people record and experience someone else’s memories.
Creating its visceral first-person sequences required nearly a year of work on a custom 35mm camera. Weighing about eight pounds, it was designed to move like the human eye. The production was also almost entirely nocturnal, with 77 of its 80 shooting days taking place at night.
The film’s volatile vision of Los Angeles was partly shaped by the 1992 riots, whose aftermath Bigelow witnessed while helping with the cleanup. James Cameron had conceived the story years earlier, while Bigelow and screenwriter developed it into the darker, more political film that reached theaters.
Critics were divided upon release. Roger Ebert gave it four stars, calling it “a technical tour de force” and praising Bigelow for using SQUID to make viewers confront the action rather than simply consume it as spectacle.
Despite its $42 million budget, the film was a commercial failure upon release, but it gained much more appreciation over the years.
We need more examples like this in the open-source RL ecosystem
Very well-written and articulated blog by @lu_jasper on training search agents with GRPO was a nice weekend read !!