So many new fun new usecases are going to be unlocked when the MoE version of 3.8-27B drops! I'm already getting more than 100 tokens/s on the int4 quant with speculative decoding. 100 tokens/s running 27/4 on this model is a lot of home-lab intelligence!
As an experiment, I tried to get KVM-accelerated macOS arm64 VM working on M2-Asahi Linux, with Graphics Acceleration powered by Reims vGPU.
It works! Source code is on Github: https://t.co/Ne8YLQvNDj
@dhh I would pay money for a clean room Markdown WYSIWYG implementation that doesn't keep breaking (either in code or in rendering) and has a decent API for highlights
Is this unmodified macOS Ghostty running on linux GPU-accelerated? Yes. Yes, it is!
✨ Mach-O Milan sneak peak ✨
This builds on top of the metal2vulkan work that's used in Reims vGPU
@rusabuilds Nothing has been easy so far :D
The challenges you're describing are real tho, they've been worked around far as I understand? Both reims-vgpu and metal2vulkan are open source! You can read on how we solved them
Reims vGPU has a discord now! As we're kicking off compatibility sprint for macOS 14+, you're welcome to join in and try out the experiments and share feedback!
https://t.co/cdWv2V6AVd
Reims vGPU Progress Update:
- Crashes are gone, haven't seen one in days
- Performance is better (but not the focus yet)
- Safari web content is flawless now
- Chrome is still faster (unlike Safari, Chrome is not using GPU acceleration)
Enjoy the video!