Homelab local AI. oMLX on Apple Silicon for real-time + code review; Ryzen AI Max+ 395 for automation, Spark for sub-agents. AI Gateway the big unlock for me.
One lesson from running local AI for real work: stop asking one machine to do every job.
Apple Silicon (oMLX) handles real-time work and code review — the stuff where I’m in the loop. A Linux box on AMD Ryzen AI Max+ 395 (Strix Halo) takes a lot of the automated workloads. Same lab. Different jobs. Fewer “why is everything stuck” evenings.
The unlock isn’t more GPUs — it’s assigning roles. Interactive + review on the box that stays snappy. Automation on the box built to grind.
Homelab local-AI series from a real multi-box setup — next posts go into the gateway and tracing layers.
@volatilemarkts@AshHart Seeing enough replies now that I'm going to have to test this on my m3 ultra 256GB. You'r not seeing ANY quality degradation at all?
I'm in that tension too. I still pay for a few subscriptions, but I've got local + OpenRouter + subs behind LiteLLM and I'm testing whether one dedicated local job can knock out a paid slot. The specific-job framing is right. Reliability on that one workflow matters more than a general vibes check.
@Ted_Cahall@ashxhart@Apple Same itch. Spark for the heavier jobs, Mac Studio for the interactive loop, and the interconnect is where it gets annoying. A ConnectX-7 port on a future Studio would make a lot of local-AI folks very happy.
@ivanfioravanti Good to see the version bump show up in real numbers. I run oMLX on an M3 Ultra for interactive work, and those Code 131 / Prose 110 peaks on 26.9.6 are solid. Version-to-version benches beat one-off screenshots for me.
That concurrency cliff matches what I see on Apple Silicon. Fine for interactive work, ugly once you stack requests. oMLX is still the engine I stick with on the Mac Studio for that reason. Going to dig into TensorFold next after seeing some Spark bake-offs. Curious how far the kernel work gets you on the multi-request side.
@WescheNex1q Same-box, same-prompt, byte-identical output is the bake-off format I actually trust. That TensorFold vs vLLM gap on a Spark is wild. I keep a Spark around for the heavier CUDA work and these are the numbers I watch before I change engines.
@ashxhart@Klarna If it comes in over $20k at reasonable config gonna have to pass as much as I’d love to have one. So glad I bought my 255GB m3 ultra before any of the price hikes.
@jun_song It's hard to argue against Mac for actual bang for the buck, and if you run a couple / few parallel requests it really is impressive what you can get out of even larger models in aggregate t/s.
@Chris_Wozniczek Haven’t run ternary Bonsai myself. Aggressive quants usually fall apart on structured output and tool calls before the chat evals look bad. Curious which of those showed up in your 20 hours.
@Oluwaphilemon1 Curious what was chewing the context by the end of those 5 hours. Agent transcripts, game assets, or both? Watching it drop from ~50 tok/s fresh to ~22 average tracks with how these long agent runs usually go.
@TeksEdge Wild how much of the progress lately is just squeezing models onto smaller boxes. Getting an image model that size down into a few GB so more people can actually run it is the kind of innovation I get geeked about.
@RussellBal 2-3x from custom kernels is a big claim in a good way. Curious whether that held on longer contexts or mostly short prompts, and if cvec actually replaced a fine-tune for you or just got you closer.
@astonepokes LiteLLM. Separate keys per provider get messy once you want fallbacks and spend caps in one place. Gateway so the app speaks one API and routing lives outside the agent code.
@pauliusztin_ Once you’re paying GPU hours instead of tokens, a tokens/sec bump is just fewer hours on the meter. That Sonnet vs self-host math makes the point pretty clean.
@DogukanUrker Those prefill numbers on a 3060 are what jumped out. Curious if the ~1600 tok/s still holds once you’re deep into that 262K window, or if it’s mostly short / cold context.
@DirtyTesLa I bought one anyway. Not because it beats a Mac Studio on every bench. I needed a third role for CUDA-shaped sub-agent work next to Apple Silicon and Strix.
@ashxhart This is useful. After your Spark posts I’ve been hunting ways to link the boxes I already have (Mac Studio + Strix, and now Spark). Thunderbolt showing up cleanly on both sides is the kind of signal I care about.