Your thoughts become your words.
Your words become your behavior.
Your behavior becomes your habits.
Your habits become your values.
Your values become your destiny.
-Mahatma Gandhi
The ridiculous arguments by Ambani’s army of Internet shills make no sense
On the one hand, they say Starlink couldn’t possibly compete with the amazing low prices offered by Ambani’s de facto monopoly in India and then, in the same breath, they essentially claim that Starlink would be too strong of a competitor! This is absurd hypocrisy and hyperbole.
The reality is that competition would greatly benefit the average citizen in India by giving them more choices, as it always does. Ambani’s profits will be lower, but he is not going out of business. And, obviously, if Starlink is too expensive in India, no one will buy it, so SpaceX will need to make pricing affordable.
Jean Drèze, one of the world's great living development economists and one of India's most remarkable public figures, picked up by the Delhi Police for simply speaking in protest.
The times we live in.
If you learn AI Infrastructure Engineering, you will never be unemployed again.
Stage 1: Systems & GPU Foundations
- Learn: C++, Rust, CUDA memory hierarchy, PCIe vs NVLink, thread blocks and warp scheduling.
- Practice: Write a custom CUDA kernel for matrix multiplication from scratch and profile it against cuBLAS.
- Why: You cannot optimize what you do not understand at the silicon level.
Stage 2: Transformer Inference Physics
- Learn: Prefill vs Decode phases, arithmetic intensity, memory bandwidth bottlenecks, KV-cache math.
- Practice: Profile a HuggingFace model with Nsight Systems to identify the exact memory bottleneck during token generation.
- Why: LLM inference is almost always memory-bound, not compute-bound.
Stage 3: Modern Serving Engines & Batching
- Learn: vLLM, SGLang, PagedAttention, continuous batching, chunked prefill.
- Practice: Deploy a 70B model and tune chunk sizes to maximize throughput without starving decode requests during long-context prefills.
- Why: Naive batching leaves 60% of GPU VRAM wasted. PagedAttention fixes this.
Stage 4: KV-Cache & Memory Optimization
- Learn: Prefix caching, KV quantization, CPU offloading, multi-turn cache reuse.
- Practice: Build a routing proxy that directs requests with identical system prompts to the same replica to share KV blocks.
- Why: Reused cache is free speed. It can slash Time-To-First-Token (TTFT) by 80%.
Stage 5: Quantization & Compression
- Learn: FP8, INT4, AWQ, GPTQ, Sparsity, TensorRT-LLM calibration.
- Practice: Serve a model in FP8 vs FP16 and benchmark the exact perplexity drop against the latency and VRAM gains.
- Why: Quantization is the only way to fit frontier models on edge GPUs and protect your margins.
Stage 6: Kernel-Level Engineering
- Learn: Triton, FlashAttention, CUDA graphs, operator fusion.
- Practice: Write a fused Triton kernel for RMSNorm or Softmax and benchmark it against the native PyTorch implementation.
- Why: Python overhead kills inference. Fused kernels save milliseconds that compound at scale.
Stage 7: Distributed Inference & Parallelism
- Learn: Tensor parallelism, Pipeline parallelism, Expert parallelism (for MoE models).
- Practice: Shard a 405B parameter model across 8 nodes and measure the communication overhead between GPUs.
- Why: Single-GPU inference is dead for frontier models. You must master the shard.
Stage 8: Speculative Decoding
- Learn: Draft-target models, Medusa heads, acceptance rates, n-gram drafting.
- Practice: Build a pipeline where a local 3B model drafts tokens for a cloud 70B model to verify in parallel.
- Why: 2x decode speed at zero quality cost is the closest thing to a free lunch in inference.
Stage 9: Multi-Node & Hardware Interconnects
- Learn: NCCL, RDMA, InfiniBand, NVLink, Disaggregated Prefill/Decode architectures.
- Practice: Set up a multi-node cluster and profile the network latency of tensor parallelism across nodes vs within a single node.
- Why: Network latency is the new GPU bottleneck. Disaggregating prefill and decode is the 2026 meta.
Stage 10: Cluster Orchestration & GPU Scheduling
- Learn: Kubernetes GPU operators, Ray, Slurm, MIG (Multi-Instance GPU) partitioning, KEDA.
- Practice: Build a queue-based autoscaler that spins up spot GPUs based on pending inference requests and drains them when empty.
- Why: Idle H100s burn $3+/hr. FinOps and scheduling are now core infra responsibilities.
Stage 11: AI Gateways, Routing & Observability
- Learn: TTFT/ITL SLOs, semantic routing, DCGM metrics, OpenTelemetry for LLMs.
- Practice: Build a gateway that routes simple queries to a quantized local model and complex reasoning to a frontier API based on prompt complexity.
- Why: Routing protects your margins and DCGM metrics tell you when your GPUs are silently throttling.
Stage 12: Public Benchmarks & Teardowns
- Learn: Reproducible methodology, latency/throughput Pareto curves, cost-per-token analysis.
- Practice: Publish a teardown comparing vLLM vs SGLang vs TensorRT-LLM on your specific hardware with full configs.
- Why: Public proof of hardware mastery gets you hired instantly by top AI labs.
Wrappers are a commodity. AI Infrastructure is the physics of scale.
The modern AI Infra Engineer builds the silicon nervous system for global intelligence.
Bookmark and Repost!
Been going through this MIT course on efficient ML / inference by @songhan_mit and it’s amazing!
Covers pruning, quantization, distillation, LLM deployment, diffusion models, etc.
Fall 2026 course, so more lectures should keep getting added 📚
https://t.co/PFJtoylmo3
There are 2 career paths in AI right now:
The API Caller:
Knows how to build with LLMs.
The Architect:
Knows how LLM systems are built.
If you want to move toward the second, Stanford has one of the best free LLM engineering playlists on YouTube:
CS336: Language Modeling from Scratch.
The 2026 course has 18 lectures covering almost the entire LLM stack -
➡️ Build the model: Tokenization, Transformers, architectures, MoE
➡️ Understand the hardware: FLOPs, memory, GPUs, TPUs
➡️ Make it fast: Triton, GPU kernels, parallelism, distributed training
➡️ Train it: Scaling laws, data collection, filtering, deduplication
➡️ Run it: Inference, evaluation
➡️ Post-train it: SFT, RLHF, RLVR
Plus multimodality.
And you don’t only watch lectures.
> You implement the tokenizer, Transformer and optimizer.
> You write FlashAttention2 in Triton.
> You build memory-efficient distributed training.
> You turn raw Common Crawl dumps into pretraining data.
> You fit a scaling law.
> You use SFT + reinforcement learning to train a language model for mathematical reasoning.
Stanford says students write at least an order of magnitude more code than in most other AI classes.
Stanford CS336. Spring 2026. 18 lectures. Free on YouTube.
Choose your path.
(Playlist in the comments)
♻️ Repost to save someone $$$ and a lot of confusion.
JUST IN: President Trump has struck a deal with Russia to supply diesel fuel to US and global markets
>300,000 tons immediately
>500,000 tons in November
>1 million tons in December
>Additional 3 million tons dependent on refinery conditions
Woah.
This semester I am teaching a new course on Agentic Systems, which explores the full stack of AI systems: from building and optimizing agent harnesses to serving and optimizing LLM serving infrastructure.
Students learn by building, measuring, and optimizing real agents and systems, balancing capability, latency, and cost.
Our goal is not just to teach today's AI technologies, but to prepare students to build tomorrow's.
All lectures, readings, and assignments are publicly available at https://t.co/z69r5AwsKX
#India: We are concerned by reports of mass detentions of demonstrators, protest leaders, members of civil society organisations, lawyers and journalists, before and during today’s large-scale demonstrations in Delhi.
We call on the authorities to ensure that the right to peaceful assembly is fully respected and fulfilled.
https://t.co/tqrQyZoXBJ
It's insane how we went from learning different languages to not writing code at all, not applicable in India cuz we're often taught coding by writing stuff on paper😋
Mujhe kya it's the nature of the society to keep people dumb and uninformed hence exploitmaxxing.#Enlightenment