Stanford just released a 1.5-hour lecture on “LLM Architecture.”
This is the exact thing systems engineers at Anthropic and OpenAI require to understand at a deep level.
Give it some time.
This might be the highest-ROI learning you do this month.
Dear recruiters, if you are writing a job posting for AI Engineering, here is how long each tool has been available, so you don't make a fool of yourself:
TensorFlow: 17 years
MCP: 6 years
vLLM: 7 years
Ollama: 10 years
CrewAI: 12 years
CUDA: 25 years
JAX: 11 years
Weaviate: 14 years
HF Transformers: 19 years
Triton (OpenAI): 10 years
LlamaIndex: 8 years
LangSmith: 8 years
AutoGen: 22 years
LangGraph: 22 years
Things to research when bored:
-String theory
-Dark matter
-Analects of Confucius
-The Fermi Paradox
-Quantum Entanglement
-Time dilation and relativity
-Transhumanism
-Lost Civilisations and Myths
-Political Bias in Cartography
-Bioluminescence
-Street art movements
-Legends of Werewolves in Europe
-The Voynich Manuscript
-Green children of Woolpit
Best GitHub Repos to Build Real AI Skills in 2026:
1. OpenAI Cookbook
https://t.co/Kzczt4Ovzm
2. OpenAI Agents SDK
https://t.co/OLSaUPyNsM
3. OpenAI Evals
https://t.co/yC2ICTvkVC
4. PydanticAI
https://t.co/Dt8m4vjZMi
5. Hugging Face Agents Course
https://t.co/ZIL1W2dUoL
6. AI for Beginners
https://t.co/9Uu9IyX0my
7. Hugging Face 101 Course
https://t.co/EwxC6EFWld
8. Hugging Face Smol Course
https://t.co/C2h7aH7MlG
9. AI Engineer Handbook
https://t.co/aeujyhE0Yk
10. AI Engineering Field Guide
https://t.co/QkybsQMLpN
Bookmark this.
12 System design concepts engineers should know:
1. Load balancing algorithms explained
↳ https://t.co/VCLCKOZzni
2. gRPC clearly explained
↳ https://t.co/QwgTXr1N9z
3. How HTTPS actually works
↳ https://t.co/wc3CQOsmPS
4. Database caching strategies
↳ https://t.co/23QdZATj2o
5. System design quality attributes
↳ https://t.co/v9WJoUPevt
6. Health checks vs heartbeats
↳ https://t.co/r5SalP6CCh
7. CI/CD pipelines
↳ https://t.co/SM2YvhioIX
8. API gateway vs load balancer vs reverse proxy
↳ https://t.co/Tg3EhT60tU
9. Microservices clearly explained
↳ https://t.co/1CpY04nNxb
10. How JWT works
↳ https://t.co/Kuv7DAj6B9
11. Idempotency in API design
↳ https://t.co/2sItwlz1oe
12. API protocols made simple
↳ https://t.co/2CEu4Wnhsv
What else should make the list?
What concepts would you like me to cover?
👋 PS: Get our System Design Handbook FREE when you join our newsletter. Join 30,001+ engineers: https://t.co/8uVCeyVa1w
--
📌 Save for later.
♻��� Repost to help other engineers learn system design.
➕ Follow Nikki Siapno + turn on notifications.
It's over. Karpathy just open-sourced an autonomous AI researcher that runs 100 experiments while you sleep.
You don't write the training code anymore.
You write a prompt that tells an AI agent how to think about research.
The agent edits the code, trains a small language model for exactly five minutes, checks the score, keeps or discards the result, and loops. All night. No human in the loop.
That fixed five-minute clock is the quiet genius. No matter what the agent changes, the network size, the learning rate, the entire architecture, every run gets compared on equal footing. This turns open-ended research into a game with a clear score:
- 12 experiments per hour, ~100 overnight
- Validation loss measures how well the model predicts unseen text
- Lower score wins, everything else is fair game
The agent touches one Python file containing the full training recipe. You never open it. Instead, you program a markdown file that shapes the agent's research strategy.
Your job becomes programming the programmer, and this unlocks a strange new loop:
1. Agents run real experiments without supervision
2. Prompt quality becomes the bottleneck, not researcher hours
3. Results auto-optimize for your specific hardware
4. Anyone with one GPU can run a research lab overnight
The best AI labs won't just have the most compute.
They'll have the best instructions for agents who never sleep, never forget a failed experiment, and never stop iterating.
𝗪𝗵𝗮𝘁'𝘀 𝘁𝗵𝗲 𝗺𝗮𝗶𝗻 𝗽𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝘃𝗶𝘁𝘆 𝗸𝗶𝗹𝗹𝗲𝗿 𝗳𝗼𝗿 𝗱𝗲𝘃𝗲𝗹𝗼𝗽𝗲𝗿𝘀?
Context switching.
You send a colleague a "quick" Slack message. Takes you 𝟱 𝘀𝗲𝗰𝗼𝗻𝗱𝘀, but it costs them 𝟮𝟯 𝗺𝗶𝗻𝘂𝘁𝗲𝘀 to get back into deep focus.
Now imagine this happening 5-10 times a day. Your best engineers aren't shipping slow. We keep breaking their flow.
It takes about 𝟭𝟱 𝗺𝗶𝗻𝘂𝘁𝗲𝘀 of uninterrupted work to reach flow state. One notification can break it, even if the developer never opens it.
Here is what helps:
- 𝗗𝗲𝗳𝗮𝘂𝗹𝘁 𝘁𝗼 𝗮𝘀𝘆𝗻𝗰. Not everything needs an instant reply
- 𝗕𝗹𝗼𝗰𝗸 𝗳𝗼𝗰𝘂𝘀 𝘁𝗶𝗺𝗲. Protect 2-4 hours for deep work daily
- 𝗕𝗮𝘁𝗰𝗵 𝗾𝘂𝗲𝘀𝘁𝗶𝗼𝗻𝘀. Combine five scattered pings into one message
- 𝗗𝗼𝗰𝘂𝗺𝗲𝗻𝘁 𝗮𝗻𝘀𝘄𝗲𝗿𝘀. If it's asked twice, write it down once
One team I worked with blocked focus time for a month. Story completion went up 𝟯𝟱%. Bug reports dropped 𝟮𝟴%.
Stop treating your team's focus like it's free. It's the most expensive resource you have.
30 security rules for AI VIBE CODING :
1. Set session expiration (JWT max 7 days + refresh rotation)
2. Never use AI-built auth. Use Clerk, Supabase Auth, or Auth0
3. Never paste API keys into AI chats. Use process.env
4. .gitignore is your first file in every project, not the last
5. Rotate secrets every 90 days minimum
6. Verify every package the AI suggests actually exists before installing
7. Always ask for newer, more secure package versions
8. Run npm audit fix right after building
9. Sanitize every input. Use parameterized queries always
10. Enable Row-Level Security from day one
11. Remove all console.log statements before shipping
12. CORS should only allow your production domain. Never wildcard
13. Validate all redirect URLs against an allow-list
14. Apply auth + rate limits to every endpoint, including mobile APIs
15. Rate limit everything from day one. 100 req/hour per IP is a start
16. Password reset routes get their own strict limit (3 per email/hour)
17. Cap AI API costs in your dashboard AND in your code
18. Add DDoS protection via Cloudflare or Vercel edge config
19. Lock down storage buckets. Users should only access their own files
20. Limit upload sizes and validate file type by signature, not extension
21. Verify webhook signatures before processing any payment data
22. Use Resend or SendGrid with proper SPF/DKIM records
23. Check permissions server-side. UI-level checks are not security
24. Ask the AI to act as a security engineer and review your code
25. Ask the AI to try and hack your app. It will find things you won't
26. Log critical actions: deletions, role changes, payments, exports
27. Build a real account deletion flow. GDPR fines are not fun
28. Automate backups and test restoration. An untested backup is nothing
29. Keep test and production environments completely separate
30. Never let test webhooks touch real systems
Ship fast. But ship secure.
Most people treat CLAUDE.md like a prompt file.
That’s the mistake.
If you want Claude Code to feel like a senior engineer living inside your repo, your project needs structure.
Claude needs 4 things at all times:
• the why → what the system does
• the map → where things live
• the rules → what’s allowed / not allowed
• the workflows → how work gets done
I call this:
The Anatomy of a Claude Code Project 👇
━━━━━━━━━━━━━━━
1️⃣ CLAUDE.md = Repo Memory (keep it short)
This is the north star file.
Not a knowledge dump. Just:
• Purpose (WHY)
• Repo map (WHAT)
• Rules + commands (HOW)
If it gets too long, the model starts missing important context.
━━━━━━━━━━━━━━━
2️⃣ .claude/skills/ = Reusable Expert Modes
Stop rewriting instructions.
Turn common workflows into skills:
• code review checklist
• refactor playbook
• release procedure
• debugging flow
Result:
Consistency across sessions and teammates.
━━━━━━━━━━━━━━━
3️⃣ .claude/hooks/ = Guardrails
Models forget.
Hooks don’t.
Use them for things that must be deterministic:
• run formatter after edits
• run tests on core changes
• block unsafe directories (auth, billing, migrations)
━━━━━━━━━━━━━━━
4️⃣ docs/ = Progressive Context
Don’t bloat prompts.
Claude just needs to know where truth lives:
• architecture overview
• ADRs (engineering decisions)
• operational runbooks
━━━━━━━━━━━━━━━
5️⃣ Local CLAUDE.md for risky modules
Put small files near sharp edges:
src/auth/CLAUDE.md
src/persistence/CLAUDE.md
infra/CLAUDE.md
Now Claude sees the gotchas exactly when it works there.
━━━━━━━━━━━━━━━
Prompting is temporary.
Structure is permanent.
When your repo is organized this way, Claude stops behaving like a chatbot…
…and starts acting like a project-native engineer.
🚨 56 researchers from 32 universities just exposed the biggest lie in AI video generation.
Every company is selling you "visual quality." Prettier videos. Higher resolution. More realistic skin and lighting.
Nobody stopped to ask: can these models actually think?
A massive coalition from Berkeley, Stanford, CMU, Harvard, Oxford, Columbia, NTU, Johns Hopkins, and 24 other institutions just built the largest video reasoning test ever created to find out.
It's called VBVR. Very Big Video Reasoning.
And the results are embarrassing for the entire industry.
Here's what they did:
They built 2.015 million video samples spanning 200 reasoning tasks. To understand how absurd that scale is: every existing video reasoning dataset in the world, combined, adds up to about 12,800 samples.
VBVR is 1,000 times larger. The paper literally draws the two circles to scale. The existing datasets are a tiny dot next to VBVR. It's almost comical.
But scale isn't even the interesting part.
They didn't just throw random video clips together. They built an entire cognitive architecture grounded in 2,000 years of philosophy. Starting with Aristotle. Literally Aristotle.
Five foundational cognitive faculties that any intelligent system should have:
Spatiality: Can the model understand where things are in 3D space? Navigate a maze? Understand geometry?
Transformation: Can it simulate how objects move, rotate, and change over time? Mental rotation. Physics.
Knowledge: Does it understand causality? Communicating vessels? Gravity? The rules of the physical world?
Abstraction: Can it solve logical puzzles? Follow algorithmic reasoning? Do the visual equivalent of Raven's Matrices?
Perception: Can it detect edges, compare sizes, count objects, identify colors and patterns?
Each faculty is mapped to parameterized task generators that produce unlimited variations. A navigation task can vary grid size, obstacle placement, start position. A rotation task can vary angles, objects, complexity. This isn't a fixed test set. It's a reasoning factory.
Then they tested every major video model on the planet.
Here are the scores:
Human baseline: 97.4%
VBVR-Wan2.2 (their fine-tuned model): 68.5%
Sora 2: 54.6%
Veo 3.1: 48.0%
Runway Gen-4 Turbo: 40.3%
Wan2.2 base: 37.1%
Kling 2.6: 36.9%
LTX-2: 31.3%
CogVideoX: 27.3%
HunyuanVideo: 27.3%
Read those numbers again.
The best commercial video model in the world, Sora 2, scores 54.6%. Humans score 97.4%. That's not a gap. That's a canyon.
And these aren't subjective aesthetic ratings. Every task has a deterministic, rule-based scorer. No AI judges. No vibes. Either the ball bounced the right way or it didn't. Either the agent found the correct path or it didn't. Either the object rotated to the correct angle or it didn't. Spearman correlation with human judgments: above 0.9.
Now here's the part most people will miss:
The five cognitive capabilities don't scale together. They found deep structural dependencies between them. And the pattern mirrors what neuroscience tells us about the human brain.
Knowledge and Spatiality are strongly correlated (ρ = 0.461). This matches the hippocampal theory: the same brain region that handles spatial navigation also supports concept learning. Edward Tolman's cognitive map hypothesis from last century, now validated in AI models.
Knowledge and Perception are strongly negatively correlated (ρ = -0.757). This aligns with the "core knowledge" debate in cognitive science: are innate abilities like object permanence really knowledge, or are they perception? The models seem to suggest they're different circuits.
Abstraction is negatively correlated with almost everything else. It shows no positive correlations with any other faculty. This is consistent with the modularity of the prefrontal cortex. Abstract reasoning is its own island.
These AI models are accidentally recapitulating real structural constraints in biological intelligence. Without anyone designing them to.
It means you can't just throw more data at the problem and expect all five capabilities to improve at once. Some of them actively compete with each other.
Here's where it gets genuinely exciting:
They took the base Wan2.2 model (37.1%) and trained it on increasing amounts of VBVR data. No architectural changes. Just data.
50K samples → scores climb steadily on both in-domain and out-of-domain tasks.
200K samples → model hits 68.5% overall. 84.6% relative improvement.
300K+ samples → performance starts to plateau.
The out-of-domain score (tasks the model never saw during training) climbed from 0.329 to 0.610. That means the model learned to reason about entirely new types of problems it was never trained on. The researchers call it "early signs of emergent generalization."
But even at peak performance: a 15% gap between in-domain and out-of-domain. And nearly 30 points below humans.
The qualitative analysis reveals something fascinating. After VBVR training, the model develops what they call "controllability-first execution logic." Instead of freely rewriting entire scenes like Sora 2 sometimes does, VBVR-Wan2.2 learns to do exactly what's asked. Delete one symbol without touching the rest. Rotate an object while keeping the background stable. Move a book to a specific slot without rearranging everything.
On one task, Sora 2 deletes the target symbol and then spontaneously rearranges all remaining symbols. VBVR-Wan2.2 just deletes the one symbol. Clean. Precise. Controllable.
They even observed "rationalizing" behavior: the model modifying intermediate elements to make its transformation narrative internally consistent. Not just producing an answer, but maintaining a coherent multi-step reasoning process.
And the honest limitations: long-horizon tasks still break. The agent sometimes duplicates or flickers during navigation. Blueprint construction can produce "correct answer, wrong method" outputs.
But here's the real takeaway nobody is talking about:
The entire AI video industry has been optimizing for the wrong metric. Visual quality is a solved problem at this point. The next frontier isn't making videos look more real. It's making videos make sense. Physics. Causality. Reasoning. Controllability.
Self-driving needs models that understand physics, not aesthetics. Robotics needs models that predict object interactions. Medical imaging needs spatial reasoning in 3D.
"Looking good" was never the goal. Thinking was.
The entire suite is open-source. The dataset (2M+ samples), the benchmark toolkit with 100+ rule-based evaluators, and the fine-tuned model are all publicly available. The pipeline supports community contributions. New tasks can be submitted, reviewed, and scaled up through their distributed generation framework.
This isn't a product launch. It's the largest open research infrastructure ever built for video intelligence. From 56 researchers across 32 universities who decided that someone needed to measure what actually matters.
Andrew Ng just revealed why the AI companies throwing the most compute at the problem are going to lose.
The winner of the intelligence race won’t use the most compute.
They’ll waste the least.
Ng: “Most of your high-dimensional data lies on a lower-dimensional subspace. It’s just a fact of life.”
Here’s what that means in practice.
You have a 10,000-dimensional dataset.
Every dimension dragged through every calculation.
Every training cycle hauling dead weight the model will never use.
Ng: “You’re carrying around these 10,000-dimensional examples throughout your whole training process.”
That bloat isn’t just inefficient.
It’s a tax on every computation you run.
Memory bandwidth. Network bandwidth. Computational speed.
All of it eaten by dimensions that contribute nothing to intelligence.
They contribute noise.
The insight that separates the architects from the arms race: that 10,000-dimensional dataset is almost entirely captured by a much smaller subspace.
The signal lives in a fraction of the space you’re paying to process.
Compress it. 10,000 dimensions down to 1,000.
Ng: “You can run your learning algorithm on a much lower-dimensional set of data and it may be much more efficient.”
Same hardware. Same budget. A fraction of the friction.
Brute force is the strategy of whoever has the deepest pockets.
Compression is the strategy of whoever actually understands the problem.
The companies that master this don’t just build faster models.
They build models that find more truth in less data than anything scaling blindly ever will.
Intelligence was never about processing everything.
It���s about knowing what to cut.