You can now optimize and make any open-source LLM faster:
1. pip install llmcompressor
2. apply quantization with 1 line of code
Two benefits:
1. Your LLM will run faster during inference time.
2. You will save a ton of money on hardware
Here are a couple of examples:
• Llama 3.1 405b requires 2 8x80GB nodes. You can optimize it using LLM Compressor to run in a single 4x80GB node with 99.9% recovery. That's 400% savings!
• Llama 3.1 70b requires 2 x 80GB GPUs. After you optimize it, it can run on a single 80GB GPU.
LLM Compressor is open-source, integrates with HugginFace model repositories, and is compatible with the most popular open-source inferencing systems, such as @vllm_project and @huggingface.
Here is the repository: https://t.co/4ExI1QgU1A
We're hosting office hours on May 15th for the community to ask questions about optimized #LLM inference, get updates on the roadmap, and generally learn about the @vllm_project.
RSVP and share what topics you want @mgoin_ and @simon_mo_ to cover here: https://t.co/uCYCpHLH87
Neural Magic is expanding to GPUs!
Complementing our existing efforts with CPUs and model compression, we just launched nm-vllm, our initial community release to support GPU inference serving for LLMs. https://t.co/NeACMVhidf
Details 👇
Sparsity and quantization make Llama 2 fast on CPUs. Simplify cloud deployments or run locally with the same software package!
If you are at #NeurIPS2023, stop by Neural Magic’s booth 520 to meet our team and feel the magic.
If you are at AWS #reinvent2023, be sure to stop by AMD's booth for an awesome demo featuring Neural Magic's DeepSparse on AMD EPYC-powered Amazon EC2 M7a instances! #deepllamaracer
Eliminate the need for specialized AI skills and complexities imposed by AI hardware. As the leader of the software-delivered AI movement, Neural Magic has recommendations to make your AI deployment efforts easier. Hear more from @addvin 👇
One of our very own Neural Magicians, @Quantum_Stat has amassed tips and tricks on how to maximize #chatgpt for #developers and #contentcreators—and we’re ready to share it with all of you🤓
Check out the only ChatGPT cheat sheet you’ll ever need: https://t.co/y4C7LgtwxK
🏁 Accelerate #YOLOv8 with Neural Magic's DeepSparse by 10x!
Developed by our partner @ultralytics, YOLOv8 takes #objectdetection to the next level with its anchor-free design.
https://t.co/6hcShPizDY
Bringing speed to object detection, while optimizing and simplifying your #YOLOv5 deployment🔥
We’re thrilled to announce our collaboration with @ultralytics, read more about the partnership on the blog🧑💻
https://t.co/JufVMqREbj
Our 2022 #YearInReview is live. We released version 1.0, continued pushing the boundaries of sparsity, and furthered incredible integration efforts with @huggingface, @awscloud, @googlecloud, @weaviate_io, just to name a few😉
https://t.co/ErTKQWzxqQ
Thank you @NEA, @a16z, @Amdocs, @ComcastVentures, @pillar_vc, and Ridgeline Partners for your continued support. And shoutout to our employees for making it all happen.
We've been working hard to make it easy to run neural networks in production without damaging your wallet or the environment. Check out our latest SOTA results for #imageclassification and how much #sparsification (pruning plus quantization) can help: https://t.co/hNdfVQKCmT
Graphic cards have long been the chip of choice for performing artificial intelligence tasks. Startup @neuralmagic wants to change that. https://t.co/WB1PslG1Yj
This week, we released the @neuralmagic Inference Engine & a suite of new tools that simplify #deeplearning performance. What's under the hood: https://t.co/Fi2fV5YJYK