GB300 Workstation and then I would slowly add AMD Halos nodes per qwen3.8-flash-next distilled from glm-5.2 runs on the GB300 if not something else I prefer.
Create a fleet of specialists and a teacher model constantly working on the gb300.
Although I may be super satisfied with glm-5.3-flash nvfp4 on max thinking in a lot of ways
@ivanfioravanti So i was LM Studio for the past year. I've played with oMLX and MTPLX and since I haven't moved from qwen3.6-35b-a3b I haven't felt the need to swap it up.
Who has feedback on vLLM-MLX?
@TechMDAI Itโs a solid deal for sure and wish I could secure it but I need a couple of things to fall in place.
Iโve seen quotes for the Dell GB300 at $180k. Donโt get that.
@Priyannkaaaa Iโve been using https://t.co/gmanBrVRn0 coding plan for over a year and have been very happy with their work. I would like to try all of these models but I would prefer to support https://t.co/gmanBrVRn0
Very surprised to see these results and I hate to ask, but I wonder if thereโs a shift with 2 x M5 Uktras vs 4 x DGX Sparks. @nvidia is clearly in the lead still.
My M3 Ultra 256GB - This is development that makes me excited for NOT selling in the aftermarket. Clustering over the years, if scalable will be great for local AI deployment and development.
I have finally decided to rent 8 c RTX pro 6000s to distill glm-5.2 in to some fine tuning training data for qwen3.5-122b-a10b and some smaller models.
So thatโs next and following should clear the path for launch. All profits from what I drop will be going towards a Supermicro GB300 workstations and smaller ai inference nodes. (GB10s, older Mac Ultras)