I really love this video and have been watching it all day. Astra is another step toward shaping the future of humanity. Welcome to the computer-use agent era.
🤔 Knowledge stored in multimodal LLM weights can be inherently limited. How can we empower multimodal LLMs with multimodal web search?
💡 In DeepMMSearch-R1, we aim to train a multimodal search agent capable of performing on-demand, multi-turn web searches and dynamically crafting queries for both image and text search tools. It can also initiate web searches based on relevant crops of the input image via using Grounding DINO as a tool.
🔗 arXiv: https://t.co/NhOPiyX4Ql
1️⃣ In SFT, we teach the model when to search, what to search for, which search tool to use and how to reason over the retrieved information.
2️⃣ SFT enables tool use, while RL refines the tool-selection behavior by reducing unnecessary calls.
Led by our awesome intern @KartikNarayan10, @tiancao, @yinfeiy etc.
Struggling with understanding–generation conflicts in unified multimodal models?
Check out Manzano 🍎🌲 from our AFM team — a simple, scalable framework with minimal task conflict, promising scaling, and SOTA results.
Excited to share Manzano from AFM team—a simple, scalable unified multimodal model for understanding and generation.
Manzano shows minimal task conflict, promising scaling behavior and state-of-the-art results among unified models.
Paper link: https://t.co/HpziryrvSc
Just read the Manzano paper. They really wrote it in Pages, Sans font, plots in Numbers O.ô
But after getting past that superficial weirdness, it's nice work. I like this figure, which clearly illustrates "big model smell"
I also like it when Parti did that already, see below.
Checked out Apple’s latest Flagship project! While focused on AToken, I admired the compute and resources on this flagship project, and the result is also amazing🎉
The legendary Ross Girshick just posted his CVPR workshop slides about the 1.5 decades he spent ~solving object detection as it relates to the ongoing LLM singularity. Excellent read, highly recommended. https://t.co/i3qS4iheU5
Language models can guide robots in multi-step visual navigation tasks like doing laundry, but current approaches require costly data sampling and annotations, which are often hard to come by.
MIT CSAIL researchers have developed LangNav, an LLM-based method that performs multi-step navigation end-to-end via textual descriptions of the scene. LangNav is more data efficient compared to VL models: https://t.co/Co8VSJlQ0e
And thank you @_akhaliq for featuring our work!
https://t.co/zoeMaeBUB1
Huge shout out to my collaborators: @rpanda89, SouYoung, Rogerio, Aude, @phillip_isola, and Yoon!
LangNav: Language as a Perceptual Representation for Navigation
paper page: https://t.co/8aBd9VnF3U
explore the use of language as a perceptual representation for vision-and-language navigation. Our approach uses off-the-shelf vision systems (for image captioning and object detection) to convert an agent's egocentric panoramic view at each time step into natural language descriptions. We then finetune a pretrained language model to select an action, based on the current view and the trajectory history, that would best fulfill the navigation instructions. In contrast to the standard setup which adapts a pretrained language model to work directly with continuous visual features from pretrained vision models, our approach instead uses (discrete) language as the perceptual representation. We explore two use cases of our language-based navigation (LangNav) approach on the R2R vision-and-language navigation benchmark: generating synthetic trajectories from a prompted large language model (GPT-4) with which to finetune a smaller language model; and sim-to-real transfer where we transfer a policy learned on a simulated environment (ALFRED) to a real-world environment (R2R). Our approach is found to improve upon strong baselines that rely on visual features in settings where only a few gold trajectories (10-100) are available, demonstrating the potential of using language as a perceptual representation for navigation tasks.
Introducing LangNav, an LLM-based navigation agent that executes multi-step navigation via textual descriptions of the visual scene. This language-based perceptual representation allows LangNav to be more data-efficient compared to vision-language models.
https://t.co/ZleSpLy9fI
Our paper has been accepted to NAACL 2024 as the findings paper!
arXiv: https://t.co/7XkYLpEUMB
GitHub: https://t.co/NCy72l0ew0
Thanks to MIT News for covering our work:
https://t.co/TPBP3qUXhW
Introducing Samba 3.8B, a simple Mamba+Sliding Window Attention architecture that outperforms Phi3-mini on major benchmarks (e.g., MMLU, GSM8K and HumanEval) by a large margin.😮 And it has an infinite context length with linear complexity.🤯
Paper: https://t.co/KwMpeyaDxc
(1/6)
A very interesting paper for Boosting Efficiency in Large Language Models: "Dense Training, Sparse Inference: Rethinking Training of Mixture-of-Experts"
Mixture-of-Experts (MoE) models offer a solution by activating only a subset of parameters, leading to improved efficiency. However, traditional MoE models often require significantly more parameters to achieve performance comparable to dense models. This is where the innovative Dense Training, Sparse Inference (DS-MoE) approach comes into play.
The Core Idea: Balancing Efficiency and Performance
DS-MoE addresses the parameter inefficiency of traditional MoE models by employing a hybrid training and inference strategy. During training, all experts in each layer are activated, enabling dense gradient propagation and efficient optimization. This ensures the model learns effectively and achieves performance comparable to dense models. During inference, however, only a subset of the most relevant experts are activated, significantly reducing computational cost and memory usage.
Key Technical Concepts
Let's delve deeper into the technical aspects of DS-MoE:
Dense Training: Unlike traditional MoEs that sparsely activate experts during training, DS-MoE computes and retains the output of all experts for gradient propagation. This enables the router to learn from the full context and optimize all experts efficiently.
Sparse Inference: During inference, DS-MoE selectively activates the top K experts based on their router scores. This K value can be fixed or dynamically determined based on a predefined threshold. The result is a significant reduction in active parameters and computational cost without sacrificing performance.
Mutual Information (MI) Loss: To ensure balanced expert utilization and prevent underfitting, DS-MoE incorporates an MI loss. This loss maximizes the entropy of the expert distribution, promoting load balancing, while minimizing the conditional entropy to ensure expert concentration and avoid overly simplistic solutions.
Mixture of Attention Head (MoA): Instead of a dense self-attention layer, DS-MoE uses a MoA layer. Each expert in the MoA layer computes query vectors, while key and value pairs are shared among all experts. This further enhances computational efficiency during inference.
Experimental Evidence and Advantages
The researchers conducted extensive experiments comparing DS-MoE with dense models and traditional sparsely trained MoE models. The results demonstrate several key advantages:
Enhanced Parameter Efficiency: DS-MoE models achieve performance comparable to dense models with significantly fewer parameters. This addresses the parameter inefficiency of traditional MoEs.
Reduced Computational Cost: During inference, DS-MoE activates only a portion of the model's parameters (30-40%), leading to significant computational savings, particularly beneficial in computation-bound scenarios like batch processing.
Improved Sparsity with Scale: Larger models exhibit greater tolerance to sparsity, effectively maintaining dense-inference performance levels with fewer active experts. This suggests the potential for even greater efficiency gains in larger-scale models.
Superior Throughput: DS-MoE excels in both computation- and I/O-bounded scenarios, offering the best throughput performance compared to both dense models and other MoE models.
Expert Sampling Strategies: Balancing Flexibility and Efficiency
DS-MoE explores different expert sampling strategies to optimize inference:
Threshold: This strategy selects experts based on their normalized probability exceeding a certain threshold. While it offers significant parameter reduction, it can be challenging for batch inference due to varying expert usage across tokens.
TopK: This approach activates a fixed number of experts, offering greater flexibility for real-world applications, particularly in batch inference scenarios.
Threshold-TopK: This hybrid method combines the strengths of both approaches. It sets a threshold for expert activation and then determines the average number of active experts per token in a batch, using this average as the K value for TopK selection.
The researchers found that all three strategies effectively balance efficiency and performance, with the Threshold strategy offering the best trade-off and the TopK and Threshold-TopK methods providing greater flexibility for real-world deployments.
DS-MoE: A Paradigm Shift in LLM Efficiency
DS-MoE represents a significant step towards achieving both computational and parameter efficiency in LLMs. By combining the benefits of dense training and sparse inference, it allows for effective learning while reducing computational cost and memory usage during deployment. This innovative approach paves the way for more efficient and accessible large language models, opening up exciting possibilities for various NLP applications.
New MoE improvement technique proposed in this paper, activating ONLY 30-40% of the model's parameters. 🔥
Runs up to 1.86× faster than similar dense models like Mistral-7B, and between 1.50× and 1.71× faster than comparable MoEs, such as DeepSeekMoE-16B and Qwen1.5-MoE-A2.7B. (using vLLM).✨
📌 Mixture-of-Experts (MoE) language models can reduce computational costs by 2-4X compared to dense models without sacrificing performance. However, MoE models generally require 2-4X times more parameters to achieve comparable performance to a dense model, which incurs larger GPU memory requirements and makes MoE models less efficient in I/O-bounded scenarios like autoregressive generation.
📌 This paper proposes a hybrid dense training and sparse inference framework for MoE models (DS-MoE) which achieves strong computation and parameter efficiency by employing dense computation across all experts during training and sparse computation during inference.
📌 DS-MoE models are more parameter-efficient than standard sparse MoEs and are on par with dense models in terms of total parameter size and performance while being computationally cheaper.
Our IBM Granite Code series models are finally released today. Despite the strong code performance that you should definitely check out, I also want to point out that the math reasoning performance of our 8B models is unexpectedly good. Congrats to all our teammates!
https://t.co/ZgWNNv1dKN
DS-MoE is both computationally and memory efficient, leading to faster processing in computation-bounded as well as I/O-bounded scenarios. The table demonstrates that DS-MoE is significantly faster than other MoEs when attempting to match the performance of 7B dense models.
Thrilled to unveil DS-MoE: a dense training and sparse inference scheme for enhanced computational and memory efficiency in your MoE models! 🚀🚀🚀
Discover more in our blog: https://t.co/XX0SHTZz3B and dive into the details with our paper: https://t.co/bCtfrbMoV8
The DS-MoE operates at the same performance level as similarly sized dense models while being much more computationally efficient. With equivalent computational cost and performance, it requires only half as many parameters compared to the Sparse MoE.