We are hiring a machine learning engineer role to drive making our research + weight releases as accessible as possible to the wider community. 🔥
If you care about model efficiency, tooling, usability, translating research into impact -- get in touch!
https://t.co/E1PG6Pruuu
New Paper Out.
Evolutionary Optimization of Model Merging Recipes
https://t.co/kn75wPgrtQ
Our goal is not about training any particular individual foundation model. Instead, we think it makes more sense to create the machinery to automatically generate foundation models for us!
Apple presents How Far Are We from Intelligent Visual Deductive Reasoning?
Vision-Language Models (VLMs) such as GPT-4V have recently demonstrated incredible strides on diverse vision language tasks. We dig into vision-based deductive reasoning, a more sophisticated but
Scaling Laws for Fine-Grained Mixture of Experts
- MoE models consistently outperform dense Transformers
- The efficiency gap between dense and MoE models widens as we scale up the model size and training budget
https://t.co/BnFe0EjgkN
New language model work! In practice, LMs often face a double constraint (i) small inference budget + (ii) little application-specific data: (i) means small specialized models for inference; (ii) means using auxiliary generic data e.g. for pretraining 1/2 https://t.co/E7MrinEcLq
Collected some of the amazing projects people are building with MLX in one place: https://t.co/iivHlPx186
Looking at that list, it's hard believe MLX is just 2 months old.
“How does such a simple objective in LLM training, next-token prediction, result in such remarkably intelligent behavior?”
This question is on everyone's mind, from everyday LLM users to expert researchers.
✨We've got a solid answer!✨
https://t.co/ujk4uYvgBj
1/7 Super excited about my Apple Internship work finally coming out:
Rephrasing the Web: A Recipe for Compute and Data-Efficient Language Modeling
TLDR: You can train 3x faster and with upto 10x lesser data with just synthetic rephrases of the web!
📝 https://t.co/1zoYmRIFhl
Apple presents Rephrasing the Web
A Recipe for Compute and Data-Efficient Language Modeling
paper page: https://t.co/ygmQs62Epa
Large language models are trained on massive scrapes of the web, which are often unstructured, noisy, and poorly phrased. Current scaling laws show that learning from such data requires an abundance of both compute and data, which grows with the size of the model being trained. This is infeasible both because of the large compute costs and duration associated with pre-training, and the impending scarcity of high-quality data on the web. In this work, we propose Web Rephrase Augmented Pre-training (WRAP) that uses an off-the-shelf instruction-tuned model prompted to paraphrase documents on the web in specific styles such as "like Wikipedia" or in "question-answer format" to jointly pre-train LLMs on real and synthetic rephrases. First, we show that using WRAP on the C4 dataset, which is naturally noisy, speeds up pre-training by sim3x. At the same pre-training compute budget, it improves perplexity by more than 10% on average across different subsets of the Pile, and improves zero-shot question answer accuracy across 13 tasks by more than 2%. Second, we investigate the impact of the re-phrasing style on the performance of the model, offering insights into how the composition of the training data can impact the performance of LLMs in OOD settings. Our gains are attributed to the fact that re-phrased synthetic data has higher utility than just real data because it (i) incorporates style diversity that closely reflects downstream evaluation style, and (ii) has higher 'quality' than web-scraped data.
ICLR24 Spotlight: To train general-purpose SSL models, it's important to measure the quality of representations during training. But how can we do this w/o downstream labels?
We propose a new label-free metric to eval SSL models, called Linear Discrimination Analysis Rank(LiDAR)
Excited to share AIM 🎯 - a set of large-scale vision models pre-trained solely using an autoregressive objective. We share the code & checkpoints of models up to 7B params, pre-trained for 1.2T patches (5B images) achieving 84% on ImageNet with a frozen trunk.
(1/n) 🧵
MLR at Apple just opened an office in Copenhagen, Denmark 🎉🎉💃🇩🇰 And we are hiring!
https://t.co/IQgjGub5m9
Please reach out if you're interested! This office opening is pretty neat just on its own, if I do say so myself. 😊 1/
Just in time for the holidays, we are releasing some new software today from Apple machine learning research.
MLX is an efficient machine learning framework specifically designed for Apple silicon (i.e. your laptop!)
Code: https://t.co/Kbis7IrP80
Docs: https://t.co/CUQb80HGut
I am looking for strong PhD interns to join Apple MLR early 2024! Topics will be around diffusion generative models broadly speaking and you’ll be in the bay area (SF/Cupertino). Apply here https://t.co/nvQDhZfeNq
Apple ML Research in Cambridge, UK is looking for a PhD intern 🎓
Topics: Self supervised representation learning, generative modeling, diffusion models.
Job Req: https://t.co/OF611EvjIC
New blog post: Collective Intelligence for Deep Learning
Recently, @yujin_tang and I published a paper about how ideas like swarm behavior, self-organization, emergence are gaining traction in deep learning.
I wrote a blog post summarizing the key ideas:
https://t.co/S644KjM20e
📢New internship opening at AI4Science in Amsterdam or Berlin!📢
For our interdisciplinary team working on electronic structure and deep learning, we are looking for someone to join us and work on synthetic data generation and curation. Please apply here! https://t.co/ov3a6C3iNK
MobileCLIP: Fast models & Fast training
🚀2.3x faster inference
🏆10% more accurate
🚀18x faster train
✂️1000x less data
Novelty🛠️:
🗝️Method: MultiModal Dataset Reinforcement
🔨DataCompDR: Improved dataset
🔧Fast image-text arch
w/ Pavan @HPouransari@litemax@OncelTuzel@Apple
Our Apple ML Research team in Barcelona is looking for a PhD intern! 🎓
Curiosity-driven research 🧠 with the goal to publish 📝
Topics: Confidence/uncertainty quantification and reliability of LLMs 🤖
Apple here: https://t.co/ZiN3ecWGo7