you can prompt this entire facility
one model controls everything: equipment, researchers, and inventory
I spent two weeks living inside it, working on C5R's launch with Astra – here's what it felt like:
In the last few months, we have been experimenting with edge inference - primarily for ASR, LLM, and TTS amongst many smaller components and agents to assist with in-cabin use cases.
The architecture choices we make have a great impact on observed latency during the voice pipelines. And each component has to work in harmony with others, find the priority requests, orchestrate which agent takes precedence, ensure the TTS is expressive, and most importantly - cohesive.
The available resources are finite and constrained, unlike cloud compute or even desktop compute. The system design and systems engineering required to piece all of these together is one of the interesting perks of my day-to-day job.
We are also playing with NVIDIA's TensorRT Edge-LLM - a reference runtime for inference. The associated comprehensive documentation provides a glimpse into the system design considerations we must look into.
Here is one such documentation on inference runtime architecture.
https://t.co/SxyvijdtVS
Went back through NVIDIA’s CUDA Refresher series this week with some new folks.
https://t.co/S5qlmh5dwY
Origins of GPU computing: why we ended up here. CPUs stopped getting faster the easy way, and parallel hardware picked up the workloads that could split.
Getting started with CUDA: what the toolkit is and how you get a first program running on the GPU.
The GPU computing ecosystem: the libraries and tools sitting on top of CUDA. Useful for knowing what you should NOT be writing yourself.
The CUDA programming model: kernels, grids, blocks, threads, how work actually maps onto the GPU.
Note: it’s from 2020, so a lot of the tooling and setup steps have moved on and may not apply anymore. But the core GPU primitives still hold up.
It's a little hard to internalize that the list of things one person can achieve is rapidly growing. You must think big to match what you're truly capable of now
i was having dinner with someone recently who asked why i spend so much time reading & posting on x. she had a pretty negative perception of the platform mostly shaped by the mainstream narrative around it.
my answer was simple. x is where the future gets beta tested.
you basically get to watch ppl build things, talk through ideas, show off stuff, & argue about where everything is going, way way before it reaches anyone else.
there really isn’t another surface on the internet quite like it.
Incredible interview of @cz_binance. He comes across as an incredibly driven and technical individual with insane levels of consistency, single minded focus and resilience.
https://t.co/yGZk2gVcJR
Claude has discovered a previously unknown enzyme system hidden in the DNA of bacteriophages. Beside the enzyme’s gene sits a long array of repeating DNA—a structure that looks somewhat similar to CRISPR.
We don’t yet understand what this system does, but only a handful of known systems share its features, and all of them are able to cut, copy, and paste DNA. Historically, the discovery of such programmable systems has helped revolutionize medicine. CRISPR, for instance, is now the foundation of genetic medicines. But it will take much more work to learn what this system does, and whether it can be put to similar use.
Read more: https://t.co/RuEosScSMb
I wrote this doc to explain KV caches and KV cache offload a while ago. I found less-technical infra folks couldn't confidently tell when users might have a problem, because all the existing online explainers are either too superficial or too math-heavy.
https://t.co/huDBgkplf9
Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family.
It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
Paper thread! I recently read this paper from some FAIR colleagues and NYU folks in depth, so making a thread while it's fresh in my head.
The main point is seeing how far on the SMALL side we can go in scaling laws without losing fit or predictability, and what it takes?
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe.
GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale.
We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
we launched the most comprehensive ai performance engineering repo in the world
now we'll be posting every single resource
this is Wafer's ai performance engineering series
save this to keep up with the series. links in thread 🧵
part 3: "Intro to CUDA C++" from NVIDIA's CUDA Programming Guide.
NVIDIA covers the execution model, memory movement, and correctness checks behind CUDA programs:
- kernel launches, grid dimensions, and the organization of threads into blocks.
- thread indexing and work assignment, including bounds checks for inputs that aren't multiples of the block size.
- unified memory and explicit memory management, including control over data placement and transfers between CPU and GPU.
- asynchronous kernel execution and synchronization before the CPU uses GPU results.
- shared memory and block-level synchronization for threads that need to exchange data and coordinate their work.
- runtime initialization and the setup costs that can affect measurements of the first runtime calls.
- error handling for kernel launches and execution, including failures that surface in later API calls.
- checking GPU results against a CPU implementation with a floating-point tolerance.
the worked examples connect these concepts in a complete vector addition program, showing how to divide the work, manage its memory, and check the results before moving on to more complex kernels.
figure from An Even Easier Introduction to CUDA (Updated)
we launched the most comprehensive ai performance engineering repo in the world
now we'll be posting every single resource
follow and save to keep up with the series. links in thread 🧵
part 4: Programming Massively Parallel Processors: A Hands-on Approach
the authors explain how GPU hardware executes parallel programs and how memory access and work distribution affect performance.
for an AI performance engineer, the book connects those fundamentals to the algorithms and optimization techniques used to diagnose slow kernels and decide what to change.
it covers:
- CUDA programming and GPU execution, including CPU and GPU cooperation, multidimensional grids, warps, scheduling, and synchronization. the vector-addition, image-processing, and matrix-multiplication examples show how to assign work to threads and reason about execution efficiency.
- memory hierarchy, coalescing, tiled matrix multiplication, thread coarsening, and occupancy. the authors explain how to reduce memory traffic while accounting for register and shared-memory usage, then provide a checklist for identifying a computation's bottleneck.
- convolution and stencils, including constant memory, caching, shared-memory tiling, and register tiling. these examples show how neighboring outputs can reuse input values, reducing repeated memory accesses when computing over arrays and grids.
- histograms, reductions, and prefix sums, including atomic operations, privatization, and work efficiency. the authors show how to reduce contention, limit divergence, and combine partial results without adding unnecessary memory traffic or computation.
- merging and sorting, including input partitioning, tiled merge, radix sort, and merge sort. these chapters connect algorithm choice and thread-to-data mapping to memory coalescing and the distribution of work across the GPU.
- sparse matrices and graph traversal, including sparse matrix-vector multiplication and breadth-first search. the authors compare storage formats and parallelization strategies to show how memory access, control divergence, and contention affect performance.
- deep learning, including perceptron inference and backpropagation, convolutional neural networks, a CUDA convolutional-layer inference kernel, convolution expressed as matrix multiplication, and cuDNN. this chapter connects the book's GPU programming techniques to the implementation of neural-network computations.
- case studies in MRI reconstruction and electrostatic potential mapping. the authors work through parallelism, loop transformations, memory layout, and validation, including how scatter and gather approaches change the cost of a computation.
- computational thinking and parallel algorithm design, including algorithm selection and problem decomposition. these provide a method for finding parallel work and choosing an implementation around the computation's requirements.
- CUDA streams and heterogeneous clusters, including MPI communication and CUDA-aware MPI. a distributed stencil example shows how to overlap communication with computation and coordinate work across GPUs.
- dynamic parallelism and advanced CUDA practices, including GPU-launched kernels, zero-copy memory, unified memory, and profiling and debugging tools. the book examines how kernels launch work and access data, including the limitations that can affect execution efficiency.
- numerical considerations, including floating-point representation, rounding, arithmetic accuracy, and numerical stability. the appendix explains how arithmetic and algorithm choices affect the reliability of computed results.
the book develops these ideas through worked kernels and applications, then connects them to neural-network computation in its deep-learning chapter.
for an AI performance engineer, that makes the material useful for reasoning about kernel execution, reducing memory traffic, and checking numerical results.
Wafer just dropped an AI performance engineering repo that covers:
> GPU fundamentals & CUDA
> kernel optimization
> FlashAttention
> KV caching, quantization
> NVIDIA, AMD & TPU architectures
and all resources link to solid docs, papers, and repos.
Build your own harness, folks.
This is absolute banger paper from NVIDIA on self-evolving agent harnesses.
(bookmark it)
They introduce SoL-Pi which cuts token traffic by nearly half.
And it matches its baseline harness on GPT-5.6 Sol and Opus 5.
More details below:
Instead of tuning a harness by hand, they run auto-research loops at the harness layer across many repository-derived and verifier-driven environments, keeping only the mechanisms that survive selection.
Four mechanisms survived:
> Action Fusion changes how actions execute
> Online Context Compact handles compaction during a run
> ObservationPack reshapes observation handling
> Evidence-Preserving Reducer covers delegated reading
On the 51-task EdgeBench evaluation, the savings translate to about a third off API cost. In dollars that is an estimated $8.75 to $13.50 per hour against native Codex and Claude Code harnesses, and $4.36 to $5.71 against the baseline harness.
Because the search runs across many environments rather than one, the retained mechanisms keep working outside the setting that produced them. Code is on GitHub under NVlabs.
Paper: https://t.co/1x26LzuE6d
Chat with Paper: https://t.co/kygTc5XLFB
RLM (recursive language model) "framework" is pretty cool, it's an LM (harness) that can programmatically manipulate its input as external data and recursively invoke itself (or other LMs) on selected portions of it
the key difference from a "conventional" coding agent is that the entire input / intermediate results can live in an external Python REPL, while the root LM decides what computations to perform (mitigating the context rot)
to make it less abstract, let's say we have the following user query:
"among questions associated with users 123 and 456, how many should be classified as 'entity' questions?"
(an entity question is something like "who is Albert Einstein?")
let's say that the full dataset contains 5k entries:
Date: ... || User: 789 || Instance: How do I bake bread? Date: ... || User: 456 || Instance: What is the capital of France? ...
important bit here: the root LM does not receive these 5k entries in its prompt, the RLM harness stores them in a var called `context`
steps in the RLM framework:
1. the root LM sees the user's query and knows that context exists (through system prompt), it then generates Python code:
print(context[:2000])
the REPL executes the code and returns the first 2k chars (this partial observation is fed into the context)
2. the root LM now understands the dataset's format. next up it filters, e.g.:
lines = context.splitlines()
relevant = [
line for line in lines
if "User: 123 ||" in line
or "User: 456 ||" in line
]
print(len(relevant))
let's say that returns 347, the root has reduced 5k entries to 347 (without polluting the context with irrelevant records)
3. next up the model can chunk up the relevant lines and recursively call itself:
chunks = [
relevant[i:i+50]
for i in range(0, len(relevant), 50)
]
results = []
for chunk in chunks:
result = llm_query(
"Classify each question as entity or non-entity. "
"Return the number of entity questions.\n"
+ "\n".join(chunk)
)
results.append(int(result))
the harness executes seven sub LLM calls in this example, and the results are now in `results` var, e.g. [17, 21, 19, 23, 20, 18, 17]
finally the model can return sum(results) as the result
this is in stark contrast compared to the context-rot that would have happened without the RLM framework
furthermore you can RL post-train the root LM inside this harness, they demonstrate much better generalization because at this level of abstraction many problems look the same - this is probably the most important bit
work by @a1zhang and @lateinteraction !
you'll know more about CUDA than 90% of people if you fully understand this guide
this is only the third resource in the ai performance engineering repo btw.
imagine the ball knowledge in the other ones
Inference scaling part 1.
Starting with a modded text generation function (temperature scaling, top-p filtering, multinomial sampling) to generate diverse outputs for self-consistency and best-of-N (improving answer accuracy by>2x)
00:00 Introduction and recap
00:31 Training-time and inference-time scaling
07:52 What we'll implement
11:47 Notebook setup and model loading
17:43 Building a flexible text generation function
24:40 Chain-of-thought prompting
28:26 Sampling and output diversity
33:43 Next-token logits and greedy decoding
38:20 Temperature scaling step by step
42:46 Softmax and token probabilities
47:42 Multinomial sampling
54:51 Adding temperature sampling to text generation
59:31 Top-p filtering step by step
1:10:23 Adding top-p filtering to text generation
1:13:43 Sampling and LLM watermarking
1:16:01 Self-consistency and majority voting
1:20:36 Implementing self-consistency
1:29:02 MATH-500 results
1:35:01 Accuracy and compute tradeoffs
1:36:50 Next steps and self-refinement