I'll be in San Diego for #NeurIPS2025
If you’re interested in edge (AR glasses, Mobile device) friendly KV cache management for streaming video, please come by and say hi!
Poster : Thu, Dec 4 (11am–2pm PST), Exhibit Hall C,D,E #4817
Paper: https://t.co/ORGHQgOzrE
Thanks to AK for sharing our paper! 🙏
We introduce EpiCache, a training-free KV cache management method that enables multi-turn dialogue on memory-constrained devices. 📱
Done as an @Apple intern, with gratitude to my collaborators!
Read more here: https://t.co/qKrsTXNOfH
How can #LLMs process/prefill infinitely long contexts on memory-constrained devices? (mobile-devices📱)
🚀 Excited to present our @Qualcomm work at #emnlp2024!
InfiniPot: Infinite Context Processing on Memory-Constrained LLMs (https://t.co/uVcZbHfTZp)
✨ Key features:
- Divide-and-conquer iterative approach with KV-Cache Compression
- Handles up to 1M tokens on NIH dataset with constrained memory (~20GB VRAM with 7B)
- +6% boost on LLaMA-3.1-8B, +4% on edge-optimized LLaMA-3.2-3B.
Come chat with us about infinite context processing!
📍 Main First Day Poster Session C (Tue 16:00 Jasmine)
(Unsolicited) Advice for a young PhD student - “This is not vocational training”
I spent most of my afternoon today at Stanford. It’s the week before autumn quarter and 8 years since I came out to the farm with wide eyes. It was as good a time as any to reflect at a place that is still the most beautiful I’ve had the chance to call home. One thing I kept thinking about was the advice I’d give to those that were now starting on this path - things have changed so much. Nevertheless, I’d reiterate the advice that had been given to me that was most meaningful in retrospect.
Before coming out west, I’d built a startup in Boston. My first boss and CEO, a professor at MIT, had finished his PhD at Stanford 9 years prior. Through the infinite wisdom of hindsight, he’d provided me many pieces of prescient advice. One that stuck out, and is somewhat under communicated, was to not treat the degrees vocationally.
A parallel he drew is the historical origin of research based academia - that of patronage towards gifted youth. You see, many of the greats, such as Gauss, were given the privilege to learn and discover the fundamental truths through noble patronage. He’d told me that framing the program in this manner would be effective at making the most of it - an exceedingly rare privilege to retreat from the world to learn, discover, meet great people, and introspect. Not as a stepping stone or accomplishment towards a set path or career outcome - a unique aspect relative to any other degree.
In this vein, the PhD is not a means to an end and extrinsically a negative EV bet (maybe this is not true anymore with 7 figure AI salaries). The way to truly win the wager is not carefully crafting the best outcome at the end of the tunnel but through intrinsic value gained in discovery and relationships - learning to be a first principles reasoner, a cogent communicator of complex thought, an adept meta learner, and building authentic connections with those who are invariably the brightest in the world. The best way to do this is to be open-minded, non-transactional and throw out any plans you made on the way in. You have a lot of time, so try a lot of things.
If you play your cards right, you get the chance to live a very interesting life, compound knowledge tremendously, and enjoy the company of the best people along the way
How should we set LoRA "rank" in Quantization-aware PEFT? (e.g., QLoRA)
- 💡2-bit Q Error shows intrinsically higher rank than 3/4-bit, but interestingly, not every sublayer requires it.
- 🛠️ RA-LoRA dynamically adjusts ranks based on Q-error characteristics of each sublayer.
RA-LoRA: Rank-Adaptive PEFT for 2-bit Quantized LLMs
https://t.co/ESqostsmmY…
Please visit my findings poster session 4 - 14 Aug 12:15 - 13:15
#ACL2024 #aclmeeting
@HanGuo97 Thank you for your outstanding work! Noticed that Equation 2 in the paper similar to the LoftQ initialization algorithm, yet order of SVD init step and quantize step is opposite. I'd be interested to hear any insights on this distinction!
Quantizable Transformers: Removing Outliers by Helping Attention Heads Do Nothing
paper page: https://t.co/zjeH0y3LTz
Transformer models have been widely adopted in various domains over the last years, and especially large language models have advanced the field of AI significantly. Due to their size, the capability of these networks has increased tremendously, but this has come at the cost of a significant increase in necessary compute. Quantization is one of the most effective ways to reduce the computational time and memory consumption of neural networks. Many studies have shown, however, that modern transformer models tend to learn strong outliers in their activations, making them difficult to quantize. To retain acceptable performance, the existence of these outliers requires activations to be in higher bitwidth or the use of different numeric formats, extra fine-tuning, or other workarounds. We show that strong outliers are related to very specific behavior of attention heads that try to learn a "no-op" or just a partial update of the residual. To achieve the exact zeros needed in the attention matrix for a no-update, the input to the softmax is pushed to be larger and larger during training, causing outliers in other parts of the network. Based on these observations, we propose two simple (independent) modifications to the attention mechanism - clipped softmax and gated attention. We empirically show that models pre-trained using our methods learn significantly smaller outliers while maintaining and sometimes even improving the floating-point task performance. This enables us to quantize transformers to full INT8 quantization of the activations without any additional effort. We demonstrate the effectiveness of our methods on both language models (BERT, OPT) and vision transformers.
@MarkSchmidty Very thanks for sharing wonderful work! Quick question, Could you provide more details on the statement that 20x faster?
Curious to know how you measured the 4-bit model generation latency and was the baseline latency for fp16 measured using PyTorch implementation?
@Karttikeya_m Thanks for sharing great solutions for painful AR inference! Learned a lot.
Quick question, is this solution also compatible with other decoding strategies? (e.g., Beam Search, Top-K sampling..)
It's here–the deepest, sharpest infrared view of the universe to date: Webb's First Deep Field.
Previewed by @POTUS on July 11, it shows galaxies once invisible to us. The full set of @NASAWebb's first full-color images & data will be revealed July 12: https://t.co/63zxpNDi4I