“Attention Is All You Need” is not actually the best first paper if you’re a software engineer transitioning to AI engineering.
I made that mistake.
Before Q, K, V, attention heads and Transformers, understand what actually happens when you type:
What is the capital of India?
- How does text become tokens?
- How do tokens become vectors?- - - How does the model predict the next token?
- How does one token become the next?
Only then does Attention Is All You Need really click.
Instead of starting with random LLM papers, I want you to follow what actually happens when you give an LLM some text.
The journey:
Text
↓
Tokenization
↓
Token IDs
↓
Embeddings
↓
Positional Information
↓
Transformer
↓
Attention
↓
Feed Forward Network
↓
Logits
↓
Next-token prediction
↓
Sampling / Decoding
↓
Generated text
01. Text → Tokens
Start by understanding why LLMs don’t directly process words or characters.
Byte Pair Encoding
https://t.co/MB4X77n88l
Neural Machine Translation of Rare Words with Subword Units
https://t.co/MB4X77n88l
Then understand modern tokenizers such as BPE, WordPiece and SentencePiece.
SentencePiece
https://t.co/13IbSHvJ3I
02. Token IDs → Embeddings
A token ID is just an integer.
The model needs to turn that integer into a vector.
Distributed Representations of Words and Phrases
https://t.co/rZYjiy4O67
word2vec
https://t.co/IymK2c5h6A
Then understand embedding spaces and why semantically related tokens can have related representations.
03. Embeddings → Positional Information
Attention doesn’t inherently know whether a token came first, second or 100th.
So how does a Transformer know the order?
Attention Is All You Need
https://t.co/xxqnVUSN0u
RoFormer: Rotary Position Embedding
https://t.co/2vWYg3llWj
ALiBi
https://t.co/kheZroFuOQ
04. Positional Embeddings → Transformer
Now understand the Transformer itself.
Attention Is All You Need
https://t.co/xxqnVUSN0u
The Illustrated Transformer
https://t.co/Ohx9ulvn3t
Don’t just memorize the architecture.
Understand the data flowing through every layer.
05. Transformer → Attention
Now Q, K and V finally make sense.
Read:
Attention Is All You Need
https://t.co/xxqnVUSN0u
Fast Transformer Decoding: One Write-Head is All You Need
https://t.co/AzHEapHUza
GQA
https://t.co/KBaiaVr1Wy
FlashAttention
https://t.co/rP7mZUuRpu
FlashAttention-2
https://t.co/0Tca2WgF9f
At this point you should understand:
Q → What am I looking for?
K → What information do I contain?
V → What information should I provide?
06. Attention → Feed-Forward Network
Attention isn’t the entire Transformer.
Every Transformer block also contains a feed-forward network.
A useful next step is understanding why these layers exist and how they transform representations.
Switch Transformers
https://t.co/WLRTvCGXTT
This also naturally leads into:
Dense models → Mixture of Experts → Sparse computation
07. Transformer → Logits
Eventually the final hidden representation has to become scores over the vocabulary.
Those scores are called logits.
Language Models are Few-Shot Learners
https://t.co/VnvqFkLGjc
This is where the architecture connects back to the original goal:
predict the next token.
08. Logits → Next Token
Now understand how those scores become probabilities.
Softmax.
Then understand why the highest-probability token isn’t always selected.
Temperature
Top-k
Top-p
Greedy decoding
Beam search
A good starting paper:
The Curious Case of Neural Text Degeneration
https://t.co/dbNysJG9RI
09. Next Token → Generated Text
The model predicts one token.
Then that token becomes part of the input.
Predict again.
And again.
And again.
Until: P(tokenₙ | token₁ ... tokenₙ₋₁)
becomes a complete sequence.
This is autoregressive generation.
GPT-3
https://t.co/VnvqFkLGjc
At this point, you should have a mental model of what actually happens when you type something into an LLM.
That’s the end of phase 1.
Thanks for reading. 🚀
Q: How can I efficiently sample production traces for review?
A: Start with broad exploration, then use signals to find likely failures. Keep some random traces in every batch so you can discover new problems.
https://t.co/vADIi01cEK
I recommend grabbing these skills that @sh_reya and I created and at the very least doing an /eval-audit of your existing pipeline.
We've found that people often find low hanging fruits! https://t.co/AkSq3NqtiJ
My second brain article passed 8 million views. So I put the whole setup, the tools, the learning material and everything around it in one place. Free
A repo and a site. What is in there:
> 109 pages on the site
> The guide: 10 sections and 65 pages, concept through troubleshooting
> 5 tracks on top of it, 44 pages, 15 of them build guides with code that runs
> 18 agent skills, 72 slash commands, 6 subagents
> 5 Python scripts: graph export, link checker, vault stats, chat converter, site builder
> A starter vault template with its own CLAUDE.md
> 87 resources: 28 tools, 26 Obsidian plugins, 15 repos, 12 skills, papers and articles
There is also a track for each of these:
> Knowledge graphs
> Jev engineering
> Agent harnesses
> Loop engineering
> Eval engineering
The Second Brain guide is a good starting point, and I'd recommend that beginners start there.
If you're a more advanced LLM user, move on to the additional resources. I'm sure you'll find a lot of useful information there.
The website and repository are constantly updated. New tooling shows up every week in this space.
Repo: https://t.co/GpYVL1Ogh5
Website: https://t.co/8dSeufUJAJ
This post was a labor of love! We’ve distilled thousands of hours of work on AI Evals into a 30 min read + skills you can use to quickly uncover errors in your product.
https://t.co/xunR1BZvVJ
BTW in addition to high quality content, Lenny effectively **pays you** in AI credits to subscribe https://t.co/9zDVkzox47 🤯 It's really good!
UPDATE: Kev-0.6B, 4B, and 8B are now available. Kev is a family of small open source Jev-like decision models you can train and run yourself.
This new family is based on Qwen3 using the same LoRA + small pointer head technique as before, but scaled up.
Out of domain, on data Kev never trained on: Kev-8B 79.6%, Jev 85.7%.
• Drop-in TypeSafe System One API; their SDK works with one `base_url` change
• Kev-4B serves on a 32 GB Mac in bf16: ~300 ms for five questions, ~40 ms on an H100
• Repeated documents hit a KV cache: 2-2.5x faster
• Apache 2.0 License. Kev-4B trains in 40 minutes on one H100. Kev-8B in 83 minutes.
Code, weights, evals: https://t.co/rTmUIMvfwh
Build your own harness, folks.
This is absolute banger paper from NVIDIA on self-evolving agent harnesses.
(bookmark it)
They introduce SoL-Pi which cuts token traffic by nearly half.
And it matches its baseline harness on GPT-5.6 Sol and Opus 5.
More details below:
Instead of tuning a harness by hand, they run auto-research loops at the harness layer across many repository-derived and verifier-driven environments, keeping only the mechanisms that survive selection.
Four mechanisms survived:
> Action Fusion changes how actions execute
> Online Context Compact handles compaction during a run
> ObservationPack reshapes observation handling
> Evidence-Preserving Reducer covers delegated reading
On the 51-task EdgeBench evaluation, the savings translate to about a third off API cost. In dollars that is an estimated $8.75 to $13.50 per hour against native Codex and Claude Code harnesses, and $4.36 to $5.71 against the baseline harness.
Because the search runs across many environments rather than one, the retained mechanisms keep working outside the setting that produced them. Code is on GitHub under NVlabs.
Paper: https://t.co/1x26LzuE6d
Chat with Paper: https://t.co/kygTc5XLFB
This is f*cking gold..
Chinese researchers just dropped a paper with a brutal title: "The End of Prompt Engineering".
The argument is simple: telling agents what to do is over as a skill.
A prompt is a wish, and a wish repeated three hundred times is still a wish.
Boundaries written in a file are law -- read by every agent before every action, versioned, owned by you.
At swarm scale nothing else holds together.
Three hundred agents cannot share a vibe, but they can share a file.
So the person stops being the one writing instructions. You become the boundary architect -- you define what never happens and review the asks.
Prompts become disposable, the constraint file becomes the job.
That is exactly the system I built and documented below: one AGENTS.md, three hundred Kimi K3 readers, zero escapes in 41 days.
The whole paper in one line: write the fence before the keys ↓
Learning to fine tune is key skill of near future - newer architectures will help democratize the generative space with tokens becoming ubiquitous.
This skill when supplemented with RL techniques are critical for every engineer.
Start soon… you don’t need H100 or B200 or whatever other SOTA hardware. Start with a 3090 on a tiny models and if you have to with LORA others to start.
Learn to fine tune a model this year
Learn to fine tune a model this year
Learn to fine tune a model this year
Learn to fine tune a model this year
Learn to fine tune a model this year
Start small
Start slow
But make sure you start
@kushal_mehra is saying @sardesairajdeep’s pearls have become straws and he is having hard time clutching those. 😂😂…
That is quite a lot of intimate knowledge… nation didn’t want to know.
@Iyervval should be jealous.
cc: @shambhav15 …
😂😂😂😂😂
Dude why do you keep clutching onto straws. You're a good political analyst. I actually miss that old Rajdeep who knew the political math like the back of his hand. Hindutva is the ideological baseline of India in every age group now.
We're building the world’s largest open-access repository for documented heritage. And we are raising the capital to make this happen.
One mission: To redefine how ancient Indian manuscripts and inscriptions are preserved and accessible in the digital age. And offer true open-source DPI to scholars, institutions, and the world.
Mission: Beyond Digitization
Step 1: We are actively engineering the foundation—machine-readable knowledge graphs, Unicode-compliant cataloging, and standardized metadata schemas.
Step 2: This is where you come in.
We're looking for the visionary partners who can fund or build this with us.
Philanthropists, archival engineers, digital humanities pioneers, and tech builders.
If you're obsessed with where our history is headed, not just where it's been—we want to talk!
Open invitation: [email protected]
Support: https://t.co/8dDXvMTMiz
Whoa, Meta released a new open-weight LLM yesterday, something that hasn't happened since the good old Llama days.
Their Meta Muse Glimmer model is a 30B multimodal reasoning model with a Gemma-like architecture design. (“Glimmer” is probably a wordplay on “Spark,” the more likely capable model from which Glimmer was distilled. Muse Spark is only available through Meta’s Model API, though.)
Architecture-wise, here are some of the main points:
1. "Only" a 131k context window, compared to Qwen3.6 and Gemma 4, which support 2x that natively; it's reasonable, but maybe on the shorter end in the age of agent harnesses
2. It's a dense model, not a mixture-of-experts. (So, it's fairer to compare it to Qwen3.6 27B than Qwen3.6 30B-A3B.)
3. Hybrid attention with grouped-query attention (GQA) and sliding window attention (SWA); the SWA:GQA pattern is a 3:1 local:global ratio. Other models like Gemma 4, which uses similar components, have a 5:1 ratio for comparison.
4. It adopts gated attention for both GQA and SWA; gated attention has become quite common in recent months. It basically applies a sigmoid gate to the attention output to decide how much of the attention information enters the residual connection. The interesting point is that it uses relatively standard GQA and SWA rather than hybrid attention mechanisms such as Nemotron or Qwen3.6.
5. A very extreme GQA ratio: 32 query heads and only 2 KV heads; for comparison, Gemma 4 31B uses 32 Q / 16 KV in the local heads and 32 Q / 4 KV in the global heads. This means that Meta Glimmer has a very small KV cache.
Overall, the probably most similar architecture is Gemma 3 27B (including the Gemma-style pre/post RMSNorm placement) and Gemma 4 31B, but with some tweaks like SwiGLU instead of GeGLU activations, gated attention, and the more extreme GQA:SWA pattern mentioned before.
What stands out is its extreme KV-cache efficiency.
I.e., the KV CACHE / TOKEN ratios (in BF16) are:
- Muse Glimmer: 52 KiB (lower is better)
- Qwen3.6 27B: 64 KiB
- Gemma 4 31B: 840 KiB
Modeling-performance-wise, their own benchmarks show that it's mostly ahead of Qwen3.6. According to the independent composite benchmarks on the Artificial Analysis Intelligence Index, it's slightly behind Qwen3.6 (see figure below). So, a few days of using it will tell where it really ranks.
Overall, it looks like a solid model, particularly for agentic workflows. What stands out most is its very low memory footprint and also pretty fast prefill and decode speed. It’s also just great to see Meta releasing open weights again :).