You can now run 70B LLMs on a 4GB GPU.
AirLLM just made massive models usable on low-memory hardware.
๐ช๐ต๐ฎ๐ ๐ท๐๐๐ ๐ต๐ฎ๐ฝ๐ฝ๐ฒ๐ป๐ฒ๐ฑ
AirLLM released memory-optimized inference for large language models.
It runs 70B models on 4GB VRAM.
It can even run 405B Llama 3.1 on 8GB VRAM.
๐๐ผ๐ ๐ถ๐ ๐๐ผ๐ฟ๐ธ๐
AirLLM loads models one layer at a time.
Instead of loading everything:
โ Load a layer
โ Run computation
โ Free memory
โ Load the next layer
This keeps GPU memory usage extremely low.
๐๐ฒ๐ ๐ฑ๐ฒ๐๐ฎ๐ถ๐น๐
โข No quantization required by default
โข Optional 4-bit or 8-bit weight compression
โข Same API as Hugging Face Transformers
โข Supports CPU and GPU inference
โข Works on Linux and macOS Apple Silicon
๐ช๐ต๐ฎ๐ ๐๐ผ๐ ๐ฐ๐ฎ๐ป ๐ฑ๐ผ
โข Run Llama, Qwen, Mistral, Mixtral locally
โข Test large models without cloud GPUs
โข Prototype agents on cheap hardware
It's ready! ๐๐๐๐๐
Introducing Ellie - your email assistant! ๐
Ellie is powered by @OpenAI and will learn your writing style and reply to emails as if you wrote them ๐ฅ
If you want to be an early user, please retweet and comment below and I'll send you an invite code! ๐
Hip Hop has lost another giant
"As I walk through the valley of the shadow of death
I take a look at my life & realize there's nothing left"
Coolio ft. L.V. & Stevie Wonder! "Gangsta's Paradise" Live! [Billboard Awards 1995]
#RIPCoolio
Compute Blade development halfway, it's time to sum up some of the results. Easier to do this in the picture.
Important changes in v0.6:
- TPM 2.0 chip
- Port for @zymbit ZYMKEY4i
- USB-A for YubiKey
- Power supply up to 22W
- NVMe 22110
Thanks for your support!
#raspberrypi
As a night owl, I feel obligated to thank every establishment that operates 24/7.
So, thank you for giving me peace of mind, and not having to worry about shops closing on me.