Most Transformer interview questions sound easy until someone asks you to explain what’s actually happening under the hood.
> Why divide attention scores by √dk?
> Why can training be parallel while autoregressive decoding can’t?
> Why does GQA reduce inference cost?
> What exactly is quadratic in self-attention?
> Why does Pre-Norm train differently from Post-Norm?
I put together 25 Transformer architecture interview questions covering the mechanics behind attention, position, normalization, and sparse computation.
➡️ Attention
→ What Transformers changed vs. recurrence and convolution
→ Q/K/V and scaled dot-product attention
→ Self-attention vs. cross-attention
→ Causal masking: parallel training vs. sequential decoding
→ Multi-head attention, and why individual heads resist clean interpretation
→ Tensor shapes: dmodel, heads, dhead, dff
→ MHA vs. MQA vs. GQA and the KV-cache tradeoff
→ What actually scales quadratically
→ FlashAttention and why reducing memory movement matters
→ Sliding-window attention and local + global attention
➡️ Position
→ Why attention alone has no built-in token order
→ Absolute vs. relative positional methods
→ RoPE and why queries and keys are rotated
→ How RoPE makes attention depend on relative displacement
→ ALiBi
➡️ Representations & stability
→ Input embeddings, LM head, and weight tying
→ Residual connections
→ RMSNorm vs. LayerNorm
→ Pre-Norm vs. Post-Norm
➡️ FFNs & sparse computation
→ What the FFN contributes beyond attention
→ ReLU/GELU vs. gated FFNs such as SwiGLU
→ Why the original Transformer used a much wider FFN hidden dimension
→ Sparse Mixture-of-Experts, routing, and active vs. total parameters
I also added interviewer notes for 8 of the trickier questions: follow-up probes, common weak answers, and red flags.
Plus a section on attention sinks and a few details that are often skipped in basic Transformer explainers.
Every answer is grounded in the original papers, including work from Vaswani, Shazeer, Ainslie, Dao, Beltagy, Su, Press, Zhang, Xiong, Fedus, and Xiao.
Algorithms by Jeff Erickson - one of the best algorithm books out there.
The illustrations are simply great - I highly recommend this.
https://t.co/8G06RjGnMA
Our book "Generative AI and Stochastic Thermodynamics: A Tale of Free Energies" is out this month. With @sirui_lu97 and @wellingmax I'm posting about topics it covers. Yesterday: heat and work in variational EM.
Today: the variational free energy, and why the ELBO is one.
Measuring information content in bits is very useful. Information theory made digital communication, cryptography and machine learning possible.
But information is not just a quantity: it also has a shape. (1/6)
Every robot you see is a data firehose generating terabytes of chaos.
This hidden crisis is the #1 reason robots fail, and it's costing the industry billions.
You see hardware, but not the data swamp drowning engineers.
In 2025, a quiet revolution is fixing it. Here’s how. 🧵
curious about the training data of OpenAI's new gpt-oss models? i was too.
so i generated 10M examples from gpt-oss-20b, ran some analysis, and the results were... pretty bizarre
time for a deep dive 🧵
Got back last night from the World AI Conference in Shanghai. Megathread with photos/videos/thoughts from the conf itself + giant expo next door (ended up going back to the expo 3 times bc there were so many interesting booths)
First up: robots robots robots
(yes, inc Unitree)
Our paper was just published in the Journal of Psychiatry & Brain Science.
5 grams of creatine per day saturates your muscles, but is likely too low for the brain.
🧵1/10
Compression is the heart of intelligence
From Occam to Kolmogorov—shorter programs=smarter representations
Meet KARL: Kolmogorov-Approximating Representation Learning.
Given an image, token budget T & target quality 𝜖 —KARL finds the smallest t≤T to reconstruct it within 𝜖🧵
Excited to share our new ICML paper, with co-authors @robert_csordas and @SchmidhuberAI!
How can we tell if an LLM is actually "thinking" versus just spitting out memorized or trivial text? Can we detect when a model is doing anything interesting?
(Thread below👇)
So this blew up. Probs good to add clarification to @DimaKrotov's quote
"Computation is a physical process. We can study the flow of bits just as we study the flow of atoms"
This has always been the perspective of Hopfield Nets. Buckle up here's a primer on Associative Memory🧵
We benchmarked leading multimodal foundation models (GPT-4o, Claude 3.5 Sonnet, Gemini, Llama, etc.) on standard computer vision tasks—from segmentation to surface normal estimation—using standard datasets like COCO and ImageNet.
These models have made remarkable progress; however, it is unclear exactly where they stand in terms of understanding vision in detail. Especially when it comes to tasks beyond question-answering. How well do they understand an object's segments or geometry?
Our analyses yield an assessment that is quantitatively and qualitatively detailed and is compatible with evaluations developed in the field of computer vision over the past decades.
Observed trends:
🔹 The foundation models consistently underperform task-specific SOTA models across all tasks. However, they are respectable generalists, which is remarkable as they are presumably trained primarily on image-text-based tasks.
🔹 They perform semantic tasks notably better than geometric ones.
🔹 GPT-4o performs the best among non-reasoning models, getting the top position in 4 out of 6 tasks.
🔹 Reasoning models, e.g., o3, show improvements in geometric tasks.
🔹 The 'image generation' models, e.g., GPT-40 Image Generation, which have been natively trained multimodally, exhibit quirks. E.g., hallucinated objects, misalignment between the input and output, etc.
🔹 While the prompting techniques affect performance, better models exhibit less sensitivity to variations in prompts. We control for the variance introduced by the prompting methods in our experiments.
🌐 Detailed analyses, visualizations: https://t.co/l8OVGMaX5V
⌨️ code: https://t.co/XufFNPWndi
🧵 1/n
Computational Capability and Efficiency of Neural Networks: A Repository of Papers
I compiled a list of theoretical papers related to the computational capabilities of Transformers, recurrent networks, feedforward networks, and graph neural networks.
Link: https://t.co/JHuymYemhM
Please let me know if you have a paper on the topic, and I will add it as soon as possible.
The image was generated using ImageFX.
Never stop being a proud physicist @DimaKrotov , it's a pleasure working with you 🥳.
https://t.co/U5gimfI6B8
"Computation is a physical process. We can study the flow of bits just as we study the flow of atoms"
More juicy Dima snippets in 🧵
The world's leading AI research center completed the most comprehensive study ever on kids and AI.
They surveyed 1,800+ children, parents, and teachers in UK.
Here's what they found:
(spoiler: children are outsmarting adults on AI)