Local AI for Intel Macs and Hackintoshes with AMD GPUs: faster long-context prompts, GPU Flash Attention, improved MoE offload, multi-GPU support, local image generation, PDF OCR, vision, and more.
Fully local, no cloud, etc.
https://t.co/mhx5S8IXPy
@GodIsVoluntary@LafeLong You can try any models you like, because we need to test if everything works properly, and yes, I'm putting the turbo engine on hold because I'll be studying its implementation in the main one. Then that engine would become obsolete, and the benefit would be nil.
@GodIsVoluntary@LafeLong I made some adjustments to the kernel ops; there were some paths that still led to the CPU, but I need real-world testing. I only checked that the paths were going where they should... can you test the new version?
@GodIsVoluntary@LafeLong Yeep... it should work pretty good right now.. Some users have already reported that it works well; of course, there are some details to correct, but I've been actively fixing them.
@GodIsVoluntary@LafeLong Hey man, thats a new build for Vega II, its already working on Vegas, but it need further testing on models, check it out
https://t.co/U88WOSdUln
@GodIsVoluntary@LafeLong I'm currently waiting for a Vega Card that was donated to me, but the earthquake in my country delayed the shipment. I'm trying to sort this out blindly. I'll try to keep you updated
A DEV TURNED A USED 2019 MAC PRO INTO A 35B LOCAL LLM SERVER BY PLUGGING IN ONE AMD EGPU AND PUSHING THE RIG TO 52 TOKENS PER SECOND FOR UNDER $3,000 TOTAL
he posts a video of the mac pro under the desk, one thunderbolt cable running to a Paladin eGPU enclosure with an RX 6800 XT inside. the internal RX 6900 XT and the external 6800 XT both load the same model in parallel. the app is Tosh LLM, a free GitHub project. the developer personally fixed the multi-GPU split overnight after the user reported the bug
the same setup i recommended in my article as the $4,199 Mac Studio M3 Ultra ceiling, except built from a 2019 Mac Pro on the used market for under half the price. 35B parameters at 52 tokens per second is faster than ChatGPT Pro on a normal day, and it runs entirely on AMD GPUs Apple does not even officially support anymore
this is not a relic. it is an obsolete intel mac pro reborn as a local AI server, built on hardware Apple discontinued and software a single developer fixed yesterday by hand
Volvio el bloqueo de @fibextelecom a la red social de X, como es posible que estos momentos de emergencia aun hayan este tipo de bloqueos. Favor liberar las redes y evitar el futuro bloqueo...
@Ulul4r@Arr3ch0
Hey @fibextelecom pueden arreglar sus nodos de internet en caracas? Como es posible que para descargar una foto tarde 3 horas,si... no exagero literamente 1 horas... y tienen un jitter y latencia de 2000ms+
@GodIsVoluntary@LafeLong Hey, saw you have a Mac Pro with AMD GPUs and are struggling with local LLMs on Mac.
I made ToshLLM, a small native app that actually gets AMD cards working...
Works well on RX 68xx and testing Vega now.
Want to try? https://t.co/U88WOSesaV
Curious how it runs on yours.
No cloud. No accounts. No perβtoken cost. Everything runs on your own GPU.
Built and tuned on an RX 6700 XT (12 GB) + NootRX Hackintosh.
Code + DMG π
https://t.co/U88WOSdUln
Most localβLLM tools on macOS target Apple Silicon. Intel Macs with AMD GPUs (incl. Hackintosh) get left behind β stock llama.cpp produces *corrupted* output on AMD dGPUs.
ToshLLM fixes that. Native macOS app, open source, GPLβ3.0. π§΅