Nah, that’s not really how it plays out. GPUs from 2016 (P100s) are still doing real inference work today. Compute doesn’t just expire it depreciates in relative speed, not usefulness. Older cards get pushed down to smaller models or fine-tuning instead of scrapped. The actual risk isn’t obsolescence, it’s buying more capacity than you can use
Common LLM benchmarks, and what they actually test
✅MMLU: broad academic knowledge (57 subjects)
✅HellaSwag: commonsense reasoning
✅GSM8K: grade-school math, step-by-step
✅HumanEval: functional code correctness
✅TruthfulQA: resistance to misconceptions
✅MT-Bench: multi-turn instruction following
A high score on one doesn’t transfer to the others. Read the eval before you read the leaderboard.
Same week, two labs, two ways to make inference cheaper.
⚡️GLM-5.3-Flash reworked its attention math to shrink memory use 4.4x.
⚡️Qwen3.8-Flash just moved a 51B-param lookup table off the GPU into regular RAM.
🧠One saves compute, the other saves memory. Both cut costs ~10x.
Intel just made the most contrarian hardware bet of the year at Hot Chips:
Ditching HBM entirely.
Intel “Crescent Island” inference GPU specs:
• 480GB LPDDR5X (skipping expensive HBM)
• 350W TDP (standard air-cooling)
• 32 Xe cores
The thesis: Agentic AI isn’t compute or bandwidth-bound anymore. It’s capacity-bound.
Massive context windows and bloated KV caches make raw TFLOPs useless if your model spills off-chip.
Tokens per Watt per Dollar is the only metric that matters now.
We’re giving scientists, mathematicians, and engineers free access to our frontier models—starting with 10,000 researchers and expanding to 100,000 through 2027.
ChatGPT for Academic Researchers is built to accelerate discovery across disciplines.
Today we release LFM2.5-Encoder-230M and LFM2.5-Encoder-350M: bidirectional encoders that stay fast at long context, even on CPU.
> LFM2.5-Encoder-230M: about 3.7x faster than ModernBERT-base on CPU at 8,192 tokens. Under 30s per forward pass, versus over a minute and a half.
> LFM2.5-Encoder-350M: 4th of 14 models on GLUE, SuperGLUE, and multilingual classification, behind only three larger models, one of them nearly 10x its size.
🧵
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
did the napkin math on paper before running anything
qwen3-8b, single L40S, spec decoding, swept concurrency 1 → 256 with a paired baseline at every point
c=1 → 2.15x
c=256 → 0.53x. half the speed of just not using it
spec decode spends idle compute. at c=1 the gpu is sitting there waiting on memory so the draft's guesses cost nothing. at c=256 the gpu is already flat out, only 2.7 of every 8 guesses land, and the wasted work comes out of someone else's request
napkin said base speed 53 tok/s and saturation around batch 52. measured 46 and 64
so it's a latency tool not a throughput tool. crossover ~80 requests
147 in / 200 out. longer outputs should push that up.
@DeclanVMichaels@burkov@thinkymachines True yeah, benchmarks are basically marketing at this point , I trust nothing until it's survived a few weeks of people actually poking at it.