@iamAlexTurnbull If by “vibe coding” you mean the original Karpathy-style definition, basically prompting, accepting what the model produces, barely reading the code, and forgetting that the code even exists, then yes.
@abacaj GLM-5.3-Flash an Intelligence Index score of 57, versus 52 for GPT-5.6 Luna at max reasoning.
GLM-5.3-Flash around 49–50 output tokens/s, whereas Luna can be around 117–168 tokens/s depending on reasoning effort.
The future of content is variable-reward AI video.
Generate 5 cheap, forgettable clips.
Then spend 10x the compute on #6 and make it fucking incredible.
Nobody knows when the jackpot is coming.
So they keep watching.
We are about to automate the slot machine.
Minimax H3 Max has generates video faster than you can watch it so I hooked it to a twitch livestream! Now you can watch infinite interdimensional cable - link to the stream below
@la_bct@varunsrin It doesn't matter.
Just generate 5 cheap, sloppy videos, then spend 10x more compute on the sixth and make it insanely good.
That’s all you need.
Five pieces of slop, then a jackpot.
The most addictive content machine imaginable.
This is the beginning of the end of passive media.
Soon you won’t scroll through content made for people like you. You’ll generate an infinite world made specifically for you, in real time, and steer it with every interaction.
We are absolutely not prepared for how addictive this will be.
Okay I built it!
🍰 Infinite Slop
https://t.co/wz0a7VSBtS
An infinite and interactive AI generated live stream of slop that goes on forever and ever
Anything that you write in the chat is generated next and AI will try to connect it to the previous video so there's an actual sequence and story line
Idea by @marcantoinefon and @rehan_shei's original stream so I made it
It'd be very expensive to run this so @fal very generously sponsors this
They also fine-tuned the model that makes this possible for the first time, making Minimax H3 50x faster and as "Max" and it can now generate videos faster than you can watch them, very cool!
Let me know what you think!!!! 😊
This is the kind of thing that will make AI genuinely addictive.
Not feeds. Not TikTok. Infinite, instant video generation where you guide what happens next and the world responds to you in real time.
Even people who barely consume video or social media will struggle to resist this. Once it becomes cheap, fast and perfectly personalized, there will be nowhere to run. This stuff will come to you and eat you alive.
Directionally right but the benchmark proves a narrower point: a 128GB M5 Max with Qwen 27B can beat a congested cloud service on latency and predictability. It does not replace Claude.
Claude went 12/12, Qwen 11/12, and long-horizon agentic quality remains different. Local complements cloud.
If we see something like:
Qwen3.8-27B Q6 → 7–10+ tok/s
prompt processing → massively faster than M4
32K–64K coding contexts → comfortable
then I think €1,489 is a fantastic little local-AI machine.
If Q6 sits at ~5 tok/s and starts swapping under realistic agent workloads, I'd keep my money and use Luna.
Apple is no longer hiding it. These machines are built for local AI.
Apple unveiled the new Mac Mini with M6 their first 2nm chip. Starting at $899.
What you can run on the $899 Mac Mini (32GB) 👇
→ Qwen 3.6 27B : fits comfortably, near-GPT-4 quality
→ Qwen 3 Coder 32B : full coding assistant on your desk
→ Llama 4 8B at full precision
→ DeepSeek R1 32B for reasoning tasks
→ 13.5x faster LLM processing than M1. 4x faster than M4.
What you can run on the Mac Studio with M5 Ultra (512GB) 👇
→ Llama 4 70B at high quality quantization
→ DeepSeek V4 with hundreds of billions of parameters
→ Frontier-scale models entirely in local memory
→ Connect four Mac Studios over Thunderbolt 5 cluster them to run trillion-parameter models locally
Apple said their Mac business grew 30% last quarter largely because people are buying these to run AI models on their own desks instead of paying for cloud tokens.
The future of inference is local. Apple built the hardware for it.
The $1,499 M6 Mac mini vs GPT-5.6 Luna Max is a fascinating AI tradeoff.
Qwen3.8 27B Q6:
• ~23GB
• fully local
• private
• unlimited
• open weights
• ~6–8 tok/s expected
Luna Max:
• frontier cloud compute
• ~145 tok/s
• 1M context
• stronger coding agent
• absurdly cheap per token
The Mac probably won't beat Luna.
For the first time, owning a genuinely capable coding model for ~€1,500 feels completely reasonable.
Would you rather own the slower AI or rent the faster one?
I think the Mac mini with M6, 32GB of RAM and whatever storage is the perfect value Qwen3.8 27B machine
Full Mac, full computer, new GPU/ANE should help model perf. I don't know if you can beat this package for $1499
The M5 Max Studio tops out at 128GB not 256GB. 256/512GB are M5 Ultra configs.
RAM decides what fits, not what performs well. Bandwidth, compute, KV cache/context, architecture and quantization matter too.
Apple itself talks about clustering multiple M5 Ultra Studios for the largest frontier open-weight models.
@kimmonismus Fitting a quantized 70B model in RAM isn’t the same as matching frontier cloud AI. Neither the M6/32GB nor M5 Pro/64GB Mac mini will realistically give you GPT-5.6 Sol Ultra-level capability.