Why #google won't let me delete files to fix the issue with too much space being taken. Why can't I fix the problem you're complaining about so easily #google#googledrive#gdrive#broken#fix
Introducing Cambrian-1, a fully open project from our group at NYU. The world doesn't need another MLLM to rival GPT-4V. Cambrian is unique as a vision-centric exploration & here's why I think it's time to shift focus from scaling LLMs to enhancing visual representations.๐งต[1/n]
Meta just released 4 models today. ๐ฅ
- Meta Chameleon: 7B & 34B language models
- Meta Multi-Token Prediction LLM
- Meta JASCO: text-to-music models
- Meta AudioSeal: audio watermarking model
This is based on Meta's groundbreaking paper released back in April-2024
"Better & Faster Large Language Models via Multi-token Prediction" โจ
Original Problem it solves
Most LLMs have a simple training objective: predicting the next word. While this approach is simple and scalable, itโs also inefficient. It requires several orders of magnitude more text than what children need to learn the same degree of language fluency.
Hence, in this paper the approach is to train language models to predict multiple future words at onceโinstead of the old one-at-a-time approach.
๐ With this new approach, a 13B parameter models solves 12% more problems on HumanEval and 17 % more on MBPP than comparable next-token models.
๐ And also models trained with 4-token prediction (instead of 1) are up to 3 times faster at inference, even with large batch sizes. ๐คฏ
---
๐ Under this approach, at each position in the training corpus, we ask the model to predict the following n tokens using n independent output heads, operating on top of a shared model trunk.
๐ The proposed method uses a shared transformer trunk to produce a latent representation of the observed context, which is then fed into n independent output heads to predict the next n tokens in parallel. This factorizes the multi-token prediction cross-entropy loss into terms for each future token conditioned on the latent representation.
๐ To make the architecture memory-efficient, the forward and backward passes are carefully reorganized. After the forward pass through the shared trunk, each output head's forward and backward passes are computed sequentially, accumulating gradients at the trunk. This avoids materializing all logits and gradients simultaneously, reducing peak memory usage from O(nV + d) to O(V + d) without impacting runtime.
๐ During inference, the additional output heads can be leveraged for self-speculative decoding methods like blockwise parallel decoding and Medusa-like tree attention to speed up generation by up to 3 times, even with large batch sizes.
๐ Experiments show that multi-token prediction is increasingly useful for larger model sizes, with 13B parameter models solving 12% more problems on HumanEval and 17% more on MBPP compared to next-token models. The approach remains beneficial when training for multiple epochs.
๐ Finetuning multi-token prediction models on the challenging CodeContests dataset outperforms finetuning next-token models, demonstrating the rich representations learned during pretraining. Next-token finetuning on top of multi-token pretraining appears optimal.
๐ For natural language tasks, multi-token prediction improves performance on generative benchmarks like summarization, while not significantly regressing on standard benchmarks based on multiple choice questions and negative log-likelihoods.
๐ The authors hypothesize that multi-token prediction mitigates the distributional discrepancy between training-time teacher forcing and inference-time autoregressive generation. They provide an information-theoretic decomposition showing how multi-token prediction increases the importance of tokens relevant for the continuation of the text.
@AmericanAir oh, and you have no idea how to scale your website. It's constantly crashing and throwing 404s while I'm trying to figure out what to do with my ticket myself. #CustomerExperience#AmericanAirlines
@seattletimes Considering how much I would have paid if I was able to buy one of those properties, well I bought it for the charm and properties of the neighborhood. Most of that's already being destroyed everywhere with those idiot square houses.
INSTITUTIONAL CRISIS: Washington vowed to transform its mental health system by 2023. The state faltered. In doing so, it failed those who need care and caused a risk to public safety. https://t.co/L9OuwsJG8q
I'm really starting to dig the @mondaydotcom boards for personal projects. A little getting used to over a trello board, but all the management extras are so nice and easy to use. #productive#getitdone
Seattleโs activist mob (which is losing steam by the day) is absolutely desperate to keep people living on the streets. They are clinging to relevancy and find it in the suffering of other human beings. Sad.
@myUHC , @uhc your https://t.co/9X2nAxP5Mr website hasn't worked for months. I can't get any information on my health stats. Pretty awful service. Never had an insurance company with no way but twitter to get ahold of them.
@TeamTurboTax@turbotax since I only get answers when I make my beef public, escalations are just screwing me over and wasting my time. Each team you escalate me to can't do anything and now I have to wait for THIS escalation. 7 days of escalations. 7. Days.