Efficient Memory Management for Large Language Model Serving with PagedAttention
paper page: https://t.co/Ab4u51yZDp
High throughput serving of large language models (LLMs) requires batching sufficiently many requests at a time. However, existing systems struggle because the key-value cache (KV cache) memory for each request is huge and grows and shrinks dynamically. When managed inefficiently, this memory can be significantly wasted by fragmentation and redundant duplication, limiting the batch size. To address this problem, we propose PagedAttention, an attention algorithm inspired by the classical virtual memory and paging techniques in operating systems. On top of it, we build vLLM, an LLM serving system that achieves (1) near-zero waste in KV cache memory and (2) flexible sharing of KV cache within and across requests to further reduce memory usage. Our evaluations show that vLLM improves the throughput of popular LLMs by 2-4times with the same level of latency compared to the state-of-the-art systems, such as FasterTransformer and Orca. The improvement is more pronounced with longer sequences, larger models, and more complex decoding algorithms.
NExT-GPT: Any-to-Any Multimodal LLM
paper page: https://t.co/tNJ46Sns2e
While recently Multimodal Large Language Models (MM-LLMs) have made exciting strides, they mostly fall prey to the limitation of only input-side multimodal understanding, without the ability to produce content in multiple modalities. As we humans always perceive the world and communicate with people through various modalities, developing any-to-any MM-LLMs capable of accepting and delivering content in any modality becomes essential to human-level AI. To fill the gap, we present an end-to-end general-purpose any-to-any MM-LLM system, NExT-GPT. We connect an LLM with multimodal adaptors and different diffusion decoders, enabling NExT-GPT to perceive inputs and generate outputs in arbitrary combinations of text, images, videos, and audio. By leveraging the existing well-trained highly-performing encoders and decoders, NExT-GPT is tuned with only a small amount of parameter (1%) of certain projection layers, which not only benefits low-cost training and also facilitates convenient expansion to more potential modalities. Moreover, we introduce a modality-switching instruction tuning (MosIT) and manually curate a high-quality dataset for MosIT, based on which NExT-GPT is empowered with complex cross-modal semantic understanding and content generation. Overall, our research showcases the promising possibility of building an AI agent capable of modeling universal modalities, paving the way for more human-like AI research in the community.
Instead of collecting datasets for instruction finetuning
1) written by humans or
2) using an LLM to generate instruction-response pairs,
here's a new 3):
"Self-Alignment with Instruction Backtranslation" (https://t.co/VzgAsuOSqS), generating instructions for unlabeled text.
A Leader, The Modi, @narendramodi , who is leading India🇮🇳 from front, making it stand tall in the world and changed the way the world looked at India.
The superb words of #Elonmusk perfectly define, How Modij's leadership is winning trust worldwide.
“I am incredibly excited about the future of India. PM (Modi) really cares about India because he is pursuing us to make significant investment in India. I am a fan of Modi." @elonmusk
#ModiInUSA #ModiInUS #ModiUSVisit2023 #Elonmusk #PMModiUSVisit
@narendramodi_in@PMOIndia
Please checkout my article "Reducing Carbon Footprint in Hybrid Cloud Modernization: Importance of Sustainability and Non-Functional Requirements" https://t.co/kpnxrRCfKD
#hybridcloud#Sustainability