AMD JUST KILLED THE $200-A-MONTH AI SUBSCRIPTION.
A 65-watt box. 128GB of unified memory. 70B models running locally.
That clip is AMD's Ryzen AI Max+ 395 mini PC, held up on stage.
The whole pitch is one spec: 128GB of high-speed unified memory shared across CPU, GPU and NPU. A model that used to need a rack of VRAM just fits.
Llama 70B at full precision on the built-in GPU.
Open models up to ~100B parameters.
No cloud. No API. A mini PC pulling 65 watts.
Your prompts stay on your desk. Your data never leaves the room. Nothing metered by the token.
This runs open models - not GPT-5, not Claude Opus. For chat, coding help and agents on your own data, it's more than enough. For the absolute frontier, the cloud still wins. And the box starts around $1,500 - that's real money, but it's a few months of a $200 sub, and then it's yours forever.
The uncomfortable part isn't that AMD made a fast box.
It's that "you need the cloud for AI" was a pricing decision - and unified memory just made it optional.
Would you mass-drop subscriptions for a $1,500 box that runs your models locally, or is the cloud still worth it for what open models can't do?
The man is past 90. He's sitting at a table, drinking a Diet Coke, and casually explaining how to save a country of a billion people.
Not joking. Not philosophizing. A concrete algorithm. One sentence.
The room listens. Nobody interrupts. Because this algorithm wasn't first used on countries.
It was used on people.
1943, Alaska. He's 19. Every morning his word decides whether pilots come back alive or not. Everyone around him tried to guess the right answer. He framed a different question - and guessing became unnecessary.
Losses - zero.
Then he applied that same question to money. For seventy years straight. Warren Buffett's partner. Net worth - $2.6 billion.
In this video he says his question out loud. Twice.
Most will hear it. Almost nobody will start asking it themselves.
$1,000 into Berkshire the day Charlie Munger walked in is worth $5 million today.
He died at 99 in November 2023, worth $2.6 billion.
He built the entire method on one mental trick he learned as an Army meteorologist in 1943.
The trick is called inversion. Munger got it from a Prussian mathematician named Carl Jacobi. Instead of asking how to succeed, ask how to fail, and refuse to do those things.
Munger tested it in the war. Stationed in Nome, Alaska, his job was to keep American pilots alive.
Instead of forecasting good flying weather, he asked what would kill a lot of pilots.
Ice on the wings. Weather closed in so tight they could not land.
Rule those two out, everything else is survivable. He never lost a pilot.
He applied the same move to money for the next seventy years.
What kills a portfolio? Leverage in a drawdown. Concentration in something you do not understand. Trading against people who understand it better than you do.
Munger's entire investing career was refusing to be near any of them.
"All I want to know is where I'm going to die so I'll never go there."
He repeated that line in almost every talk for four decades. Every audience laughed. Almost nobody used it.
Retail traders still open Robinhood at 2am, add to a losing position on 5x leverage, and check the price forty-eight times before dawn.
Each of those is a way to die. Each of them is documented. Each of them is optional.
Munger is dead. Berkshire is worth nine hundred billion dollars. The trick is free.
Almost nobody will use it tomorrow.
A math professor from UCLA walked into a meeting with the most reclusive scientist at MIT. The secretary warned him: "He doesn't see anyone. If he does, you get five minutes."
Five minutes turned into nine months of work in a basement. What they built from twelve transistors now sits in the MIT Museum. And it is the reason casinos around the world rewrote the rules of roulette.
His name is Ed Thorp. He is 89. He looks 65.
In 1968 he met Warren Buffett. He drove home and said one sentence to his wife. Then he lost track of him for 14 years. When he found him again, the stock was at $982 and everyone said the opportunity was gone.
Thorp bought. Why - he explains in the interview. His logic takes 30 seconds and has nothing to do with charts.
In 1991 he was shown the portfolio of a respected financier. Thorp checked 160 trades. Within a week he knew it was fraud. He explained everything to the man who invited him - the numbers, the proof, the logic.
That man listened. Understood every word. And kept sending money into the same scheme for another 17 years.
Thorp's explanation of why a smart person with complete information makes the worst possible decision is one sentence. It is worth the entire interview.
His hedge fund ran for 20 years. Returns around 20 percent annually. Three down months across the entire run. All three under one percent.
At the peak he shut it down and refused to open another.
Every journalist asks him the same question: why. His answer is someone else's quote from a conversation between two writers at a billionaire's party. Six words after which talking about money becomes uninteresting.
The interview is 92 minutes. There are no motivational speeches. There is a man who beat casinos, Wall Street, and time - calmly explaining exactly how he did it.
A math professor from UCLA walked into a meeting with the most reclusive scientist at MIT. The secretary warned him: "He doesn't see anyone. If he does, you get five minutes."
Five minutes turned into nine months of work in a basement. What they built from twelve transistors now sits in the MIT Museum. And it is the reason casinos around the world rewrote the rules of roulette.
His name is Ed Thorp. He is 89. He looks 65.
In 1968 he met Warren Buffett. He drove home and said one sentence to his wife. Then he lost track of him for 14 years. When he found him again, the stock was at $982 and everyone said the opportunity was gone.
Thorp bought. Why - he explains in the interview. His logic takes 30 seconds and has nothing to do with charts.
In 1991 he was shown the portfolio of a respected financier. Thorp checked 160 trades. Within a week he knew it was fraud. He explained everything to the man who invited him - the numbers, the proof, the logic.
That man listened. Understood every word. And kept sending money into the same scheme for another 17 years.
Thorp's explanation of why a smart person with complete information makes the worst possible decision is one sentence. It is worth the entire interview.
His hedge fund ran for 20 years. Returns around 20 percent annually. Three down months across the entire run. All three under one percent.
At the peak he shut it down and refused to open another.
Every journalist asks him the same question: why. His answer is someone else's quote from a conversation between two writers at a billionaire's party. Six words after which talking about money becomes uninteresting.
The interview is 92 minutes. There are no motivational speeches. There is a man who beat casinos, Wall Street, and time - calmly explaining exactly how he did it.
MSI DISGUISED A $16,000 AI CLUSTER AS A DESKTOP PC.
Four computers. One case. You'd walk past it and never know.
This is MSI's EdgeXpert tower from Computex 2026. From the front - a tall workstation. Nothing special. Pop the glass side panel and you're looking at four independent machines stacked like pizza boxes inside one chassis.
Each one runs an NVIDIA GB10 Grace Blackwell superchip with 128GB of unified memory and up to 1 petaFLOP of FP4 compute. They're linked through ConnectX-7 into a high-speed switch at the back.
So yeah. Half a terabyte of AI memory standing next to your desk. No rack. No cloud. No monthly invoice.
Need to run a 70B model? One node. Fine-tuning something heavier? Two. Want to throw everything at a 400B+ monster? All four, distributed.
Before you mortgage the house: 512GB across four nodes is not 512GB unified. Every cross-node operation pays a latency penalty. And $16K is just four EdgeXperts at list price - MSI hasn't said what the tower and switch cost on top.
But here's the thing nobody's talking about. This isn't a workstation. It's a homelab rack pretending to be furniture. MSI just made "I run my own AI cluster" look like "I have a gaming PC."
Would you drop $16K+ to never rent a GPU again, or is the cloud still cheaper for what you actually run?
MSI DISGUISED A $16,000 AI CLUSTER AS A DESKTOP PC.
Four computers. One case. You'd walk past it and never know.
This is MSI's EdgeXpert tower from Computex 2026. From the front - a tall workstation. Nothing special. Pop the glass side panel and you're looking at four independent machines stacked like pizza boxes inside one chassis.
Each one runs an NVIDIA GB10 Grace Blackwell superchip with 128GB of unified memory and up to 1 petaFLOP of FP4 compute. They're linked through ConnectX-7 into a high-speed switch at the back.
So yeah. Half a terabyte of AI memory standing next to your desk. No rack. No cloud. No monthly invoice.
Need to run a 70B model? One node. Fine-tuning something heavier? Two. Want to throw everything at a 400B+ monster? All four, distributed.
Before you mortgage the house: 512GB across four nodes is not 512GB unified. Every cross-node operation pays a latency penalty. And $16K is just four EdgeXperts at list price - MSI hasn't said what the tower and switch cost on top.
But here's the thing nobody's talking about. This isn't a workstation. It's a homelab rack pretending to be furniture. MSI just made "I run my own AI cluster" look like "I have a gaming PC."
Would you drop $16K+ to never rent a GPU again, or is the cloud still cheaper for what you actually run?
A HOMELAB YOU CAN'T ACCESS REMOTELY IS JUST A SPACE HEATER.
AND OPENING A PORT TO FIX THAT IS THE WORST THING YOU CAN DO.
that clip is a homelab in a tiny rack: cheap used mini-pcs running a proxmox cluster with ollama for local ai.
but the hardware isn't the hard part. getting to it safely when you're away from home is.
the wrong way: forward a port on your router. now your services face the open internet, and every bot on the planet starts hammering the door.
the right way - and what you're seeing here: a mesh vpn like twingate, netbird or tailscale.
it creates a private encrypted tunnel between your devices and the lab. no open ports, no public ip, nothing for anyone to find.
your phone at a coffee shop talks to your proxmox box at home as if they're on the same network - and the outside world sees zero.
that's what turns a pile of mini-pcs into infrastructure you actually rely on.
no port forwarding, no exposed dashboard, no public attack surface.
the hard part isn't building the lab. it's admitting you either can't reach it - or you left the front door wide open.
A HOMELAB YOU CAN'T ACCESS REMOTELY IS JUST A SPACE HEATER.
AND OPENING A PORT TO FIX THAT IS THE WORST THING YOU CAN DO.
that clip is a homelab in a tiny rack: cheap used mini-pcs running a proxmox cluster with ollama for local ai.
but the hardware isn't the hard part. getting to it safely when you're away from home is.
the wrong way: forward a port on your router. now your services face the open internet, and every bot on the planet starts hammering the door.
the right way - and what you're seeing here: a mesh vpn like twingate, netbird or tailscale.
it creates a private encrypted tunnel between your devices and the lab. no open ports, no public ip, nothing for anyone to find.
your phone at a coffee shop talks to your proxmox box at home as if they're on the same network - and the outside world sees zero.
that's what turns a pile of mini-pcs into infrastructure you actually rely on.
no port forwarding, no exposed dashboard, no public attack surface.
the hard part isn't building the lab. it's admitting you either can't reach it - or you left the front door wide open.
THIS FANLESS BOARD IS A FULL x86 SERVER WITH A REAL PCIe SLOT - AND YOUR ENTIRE CLOUD SITS ON IT, OPEN FROM YOUR PHONE.
that clip is the zimaboard 2. a tiny fanless single-board server.
why it isn't just another raspberry pi:
it's x86, so it runs normal software - not arm-only builds
a real pcie 3.0 x4 slot: drop in a sata card, a faster nic, even a gpu
dual sata for real drives, dual 2.5gbe, up to 16gb ram
it runs zimaos, and the files app on your phone pulls up backup, photos, media, downloads - all of it served from the board sitting on your desk.
$169 to start. silent. barely draws power. runs 24/7.
point it at whatever you want: a nas, plex, pi-hole, home assistant, even a firewall or a proxmox lab.
no cloud subscription. no arm-only limits. no data parked on someone else's server.
the uncomfortable part isn't that it's small. it's that the thing you rent as "the cloud" is a fanless board you could own for what two years of storage fees cost.
THE $4,000 DGX SPARK BENCHMARKS AT 342 TOKENS PER SECOND. THEN 73.
that's not a bug. that's local ai.
the clip is a spark running llama-bench on qwen3-coder 30b next to a mac.
pp512 = 342 t/s - the spark reading your prompt.
tg128 = 73 t/s - the spark writing the answer.
same box. almost 5x gap.
prompt processing chews through context, files, long inputs. that's compute-bound. the spark belongs here.
token generation spits out text, one token at a time. that's memory bandwidth. the spark's is modest. a unified-memory mac keeps up.
so "which box is faster" has no answer.
throwing huge context at agents? spark.
long chat replies? the mac hangs.
now the uncomfortable part.
every review, every benchmark, every hype post quotes one number. "tokens per second." never says which one.
342 sells the box.
73 is what you live with.
THE $4,000 DGX SPARK BENCHMARKS AT 342 TOKENS PER SECOND. THEN 73.
that's not a bug. that's local ai.
the clip is a spark running llama-bench on qwen3-coder 30b next to a mac.
pp512 = 342 t/s - the spark reading your prompt.
tg128 = 73 t/s - the spark writing the answer.
same box. almost 5x gap.
prompt processing chews through context, files, long inputs. that's compute-bound. the spark belongs here.
token generation spits out text, one token at a time. that's memory bandwidth. the spark's is modest. a unified-memory mac keeps up.
so "which box is faster" has no answer.
throwing huge context at agents? spark.
long chat replies? the mac hangs.
now the uncomfortable part.
every review, every benchmark, every hype post quotes one number. "tokens per second." never says which one.
342 sells the box.
73 is what you live with.
YOUR "BACKUP" PSU IS COSTING YOU $37 A YEAR TO DO ABSOLUTELY NOTHING.
Two R720s. Same rack.
168W with both PSUs.
140W with one ripped out.
28 watts à 8,760 hours à $0.15 = money you're setting on fire for a power supply that's just sitting there⦠waiting.
One setting in iDRAC kills it.
Hot Spare β Not Redundant. Pull the second PSU. Done.
But here's what nobody puts in the thumbnail:
That PSU was your only failover.
One supply dies now β no handoff, no warning.
Just a dead server and a 3 AM ticket.
Caveat: 168 vs 140 came from two different machines. Different CPUs, drives, loads. Ballpark, not a lab result.
$37 a year to sleep at night.
Or $0 and a bet against Murphy's Law.
What would you pull?
A $599 APPLE BOX THE SIZE OF A SANDWICH JUST RAN DOTA 2 AND CUT 4K VIDEO FROM THE SAME DESK - AND THE TOWER UNDER YOURS SUDDENLY LOOKS RIDICULOUS.
that clip is a base mac mini m4 getting unboxed, plugged into one monitor, and handling dota 2 without breaking a sweat.
the whole machine is smaller than the speaker sitting next to it.
what's packed into the $599 box:
M4 chip, 10-core cpu and 10-core gpu
16gb of unified memory, shared directly between cpu and gpu
256gb ssd, gigabit ethernet, macos
unified memory is the whole trick. there's no separate vram to hit a ceiling on, so the gpu pulls from the same 16gb the cpu does.
that's why the box that plays war thunder at a smooth frame rate also chews through 4k video without choking.
gaming? smooth. video editing? more than handled. same little box, same afternoon.
the numbers that actually matter: $599 once. a few watts at idle. silent under the monitor.
no 500-watt tower, no discrete gpu, no upgrade cycle every two years.
the uncomfortable part isn't that a mini pc can game. it's that "you need a big tower" was sold to you for 20 years - and a sandwich-sized box just quietly broke that rule.
follow @0xAltrion for the workflows that do the half of your job you hate, and bookmark this before your next client afternoon.
A $599 APPLE BOX THE SIZE OF A SANDWICH JUST RAN DOTA 2 AND CUT 4K VIDEO FROM THE SAME DESK - AND THE TOWER UNDER YOURS SUDDENLY LOOKS RIDICULOUS.
that clip is a base mac mini m4 getting unboxed, plugged into one monitor, and handling dota 2 without breaking a sweat.
the whole machine is smaller than the speaker sitting next to it.
what's packed into the $599 box:
M4 chip, 10-core cpu and 10-core gpu
16gb of unified memory, shared directly between cpu and gpu
256gb ssd, gigabit ethernet, macos
unified memory is the whole trick. there's no separate vram to hit a ceiling on, so the gpu pulls from the same 16gb the cpu does.
that's why the box that plays war thunder at a smooth frame rate also chews through 4k video without choking.
gaming? smooth. video editing? more than handled. same little box, same afternoon.
the numbers that actually matter: $599 once. a few watts at idle. silent under the monitor.
no 500-watt tower, no discrete gpu, no upgrade cycle every two years.
the uncomfortable part isn't that a mini pc can game. it's that "you need a big tower" was sold to you for 20 years - and a sandwich-sized box just quietly broke that rule.
follow @0xAltrion for the workflows that do the half of your job you hate, and bookmark this before your next client afternoon.
YOUR "BACKUP" PSU IS COSTING YOU $37 A YEAR TO DO ABSOLUTELY NOTHING.
Two R720s. Same rack.
168W with both PSUs.
140W with one ripped out.
28 watts à 8,760 hours à $0.15 = money you're setting on fire for a power supply that's just sitting there⦠waiting.
One setting in iDRAC kills it.
Hot Spare β Not Redundant. Pull the second PSU. Done.
But here's what nobody puts in the thumbnail:
That PSU was your only failover.
One supply dies now β no handoff, no warning.
Just a dead server and a 3 AM ticket.
Caveat: 168 vs 140 came from two different machines. Different CPUs, drives, loads. Ballpark, not a lab result.
$37 a year to sleep at night.
Or $0 and a bet against Murphy's Law.
What would you pull?