We spotlight ML researchers & practitioners. High (S) fact: ~50% code contributors to ML paper implementations are practitioners collaborating with researchers
@DrJimFan There is a startup @Vapi_AI that addresses these challenges above and offers a working system with developers building on top of their open transparent platform (developers can choose the voice, AI agent etc.)
We have been running https://t.co/nPatxEP8rb service for over a year now. The number of paper authors contributing to machine learining has grown to 433k+ and code authors to 170k+. More features spotlighting researchers to come.
This approach builds on the recent method of using 3D Gaussians to model scenes, enabling the quick and accurate creation of high-quality 3D images from photos or videos, even for large and complex scenes
@xiaolonw's contributions
https://t.co/PDOHyGuGeF
3D Gaussian Splatting is great, but can it work without the pre-computed camera poses? Introducing:
COLMAP-Free 3D Gaussian Splatting
Our recent work shows not only it can, but 3D Gaussians make camera pose estimation easy (compared to NeRF) along with reconstruction.
👇🧵
It has been 10 years since the release of the word2vec paper.
Word2vec is arguably the first model (if we can even call it that, considering its simplicity—it consists of just two arrays of vectors) to demonstrate the power of distributed representation learning.
The training process is surprisingly straightforward: it involves pulling each word's vector closer to those of other words in the same sentence, while simultaneously pushing it away from a few unrelated words not present in that sentence.
By applying this process to a large corpus, we end up with vectors for semantically related words that roughly point in the same direction within a high-dimensional space.
It's worth noting that all Large Language Models (LLMs) generate context-independent vectors, similar to those learned by word2vec, as part of their training. The models use these vectors to represent any input. Subsequently, they transform these into context-dependent vectors based on the sentence.
Tomas Mikolov's contributions
https://t.co/7uD3ZEtSbh
Congratulations to Jeff Dean, Greg Corrado, & co-authors of the paper “Distributed Representations of Words and Phrases and their Compositionality”, for winning the #NeurIPS2023 Test of Time Award! This prize recognizes a highly impactful paper published at NeurIPS 10 years ago.
This paper introduces a new objective function for self-supervised learning (SSL) that leverages the geometric properties of manifolds to which input classes are mapped. The focus is on maximizing manifold capacity, which essentially measures the number of object categories that can be linearly separated. This approach aims to enhance the efficiency and effectiveness of SSL in creating representations that are more aligned with the way biological systems process visual information. The paper presents a novel model for visual representation that is competitive with current SSL frameworks and potentially more closely mimics the functioning of the primate visual system. @tedyerxa
Excited to share our results on Efficient Coding of Natural Images using Maximum Manifold Capacity Representations, a collaboration with @KuangYilun@EeroSimoncelli and @s_y_chung to be presented at #NeurIPS2023
1/n
This work shows that Large Language Models (LLMs) aligned with fine-tuning or Reinforcement Learning from Human Feedback (RLHF) are still susceptible to prompt attacks that can reveal the training data memorized by the model. It is a known fact that base models tend to memorize their training data. This study reveals that aligning models does not conceal this vulnerability. This finding holds significance, particularly for businesses that do not publicly release models trained on sensitive data but provide access to these models through an interface.
https://t.co/tOF5Z1bHR8
First author Milad Nasr's contributions
https://t.co/0F7sPb6Obt
The prompt contains repeating words - GPt3.5 turbo model's output for such a prompt revealing training data (this was confirmed by matching model output with content downloaded from internet)
Andrej Karpathy continues to be the few biological agents whose world model is worth sampling from for nuggets of insights on LLMs
https://t.co/wSU8ypJMXr
A thoughtful answer to the question - LLMs have been practically exposed to the entire corpus of human knowledge and still have not made connections that have led to a discovery, at least not yet.
Eric's contributions to research
https://t.co/cnGICVtF8H
tl;dr: Maybe learning simple things (basic knowledge, heuristics, etc) actually lowers the loss more than learning sophisticated things (algorithms associated with higher cognition that we really care about), and the sophisticated things will eventually be learned as scaling further drives down the loss.
Some rough thoughts on this issue:
As we know, LLMs are (pre)trained to minimize next-token prediction loss across human text. Ignoring fine-tuning, RLHF, etc., what they learn will be determined both by (1) what is __possible__ to learn and (2) what is __optimal__ to learn under this objective.
So what do models learn under this objective? The low-hanging fruit might be to learn token frequencies, then common bigrams, then to learn how to repeat subsequences in context (https://t.co/DDt0sY7lgq) then simple grammatical rules, then all sorts of knowledge, and so on and so on. With a small network you can only learn the low-hanging fruit (https://t.co/tsebT22ICB), but of course the *hope* with scaling is that eventually, to reduce the loss slightly more, you'll need to learn the more sophisticated algorithms for reasoning or planning or drawing connections between ideas or coming up with new explanations (https://t.co/YKN6M4D63m).
So why might current models have not learned the algorithms associated with higher cognition yet? It could be that simply (1) they are not possible for the models to learn. But one defense of scaling would be that (2) is the bottleneck:
It could be possible that the (loss improvement) ÷ (network capacity) of learning the algorithms/circuits we associate with higher cognition is still worse than for learning many additional pieces of knowledge or statistical patterns in language. Like maybe the model could have learned the "making new connections" algorithm but actually that is suboptimal and you reduce the loss more by learning lots of dumb things instead. Presumably though eventually the algorithms for reasoning, planning, making new connections, etc. would be the optimal things to learn to further reduce loss, and they would be learned with further scaling. This explanation feels slightly odd to me since I would have expected that learning the sophisticated general parts of cognition would be very frequently useful when predicting human text and therefore would have been learned by now. But maybe it's just a fact about predicting language that there is a *lot* of low-hanging fruit in learning the dumber things: https://t.co/5xykYjioeN
Explanation (1) -- there's a fundamental limitation -- is complicated. Maybe there are so many steps or so much compute needed to implement certain algorithms that it's just really hard to implement them in a forward pass of a transformer. Further scaling could maybe solve this (maybe you just needed to make the network deeper), but maybe you'd need such a massive network that it's practically impossible. Maybe it's not about expressivity but instead about whether SGD can learn the algorithms we care about and whether we have the data needed to learn them: https://t.co/EG62wIQ7L8.
So a rebuttal to your point would be about (2), but personally I don't know if that explanation is right! The bottleneck might really be about expressivity or learnability and we're just using the wrong architectures or objectives or data. Super interesting question.
Jack mentions in his tweet that his walkthrough was inspired by the walkthroughs done by @NeelNanda5. Needless to say, if this trend picks up, it would be immensely beneficial to everyone. There is nothing quite like hearing an author discuss the details, nuances, and limitations of their own work
A paper walkthrough by @jackm2003 shares an intriguing finding: grokking has been observed in non-neural architectures. Originally noted in transformers working with algorithmic datasets, grokking is a phenomenon where the accuracy on the validation set increases, but significantly after the point at which training would typically be halted due to early stopping. In early stopping, training is stopped to prevent overfitting when the model's performance on a validation set ceases to improve. However, in the case of grokking, this dramatic increase in validation accuracy occurs much later, suggesting a different dynamic at play. The authors also hypothesize that the interplay between the error and the regularization term (which they refer to as 'complexity') in the loss function might be a crucial factor in triggering grokking.
Jack Miller's contributions to research
https://t.co/JEnKB9uz3i
For anyone planning to efficiently fine tune a LLM, this article by @rasbt on LoRA could be helpful. He explains with clarity the trade-offs to consider such as choice of quantizing pretrained weights, choice of optimizers (Adam vs SGD), impact of schedulers etc.
What is LoRA? For simplicity, imagine all the weights of a model as one large matrix W. The key insight of LoRA is that, unlike in pretraining, during fine-tuning we can approximate the gradient update matrix (which is the same shape as W) with two smaller matrix thereby achieving savings in both compute and memory.
https://t.co/okWjYgn2UO
Sebastian's contributions
https://t.co/YOgyvhland
An approach for distributed training of LLM where pretraining is performed in parallel on multiple nodes and can have a large number of training steps (compared to typical Federated averaging). The model parameters are sent to a central server that finds the average change in model parameters. This is then used to update the global model using a different optimizer (Nesterov instead of AdamW used for local updates). Note also it is the model parameter and not the gradients that are used to update global model parameters. This is then dispatched to all nodes and the cycle is repeated.
Arthur Douillard @Ar_Douillard 's contributions - (first author)
https://t.co/hLniDFHr61
This paper explores a combination of approaches to filter the input to a generator to improve RAG performance
Zora's (@ZhiruoW) contribution to research. Code is also released (models used FLAN-T5 and LLAMA-2)
It suggests a combination of methods to filter irrelevant content to a generator
https://t.co/eOJ1FegpAt
Everyone is using RAG, but most of the retrieved context is noisy! 🚨
Introducing FilCo: “Learning to Filter Context for Retrieval-Augmented Generation”
TL;DR: Get rid of the irrelevant content using FilCo, and you'll get better outputs.
Preprint: https://t.co/0EwbHxVlJu
This unforgiving retrospective analysis on the anniversary of Galactica release, by its first author, @rosstaylor90, sets a high bar for researchers - it underscores the importance of releasing work for feedback and continual improvement, even if it is a work in progress. Interestingly, the initial almost vitriolic criticism of Galactica for its creative responses is now largely accepted, viewed as either an opportunity or a limitation depending on the use case. More significantly, the limitations of base LLMs, like Galactica, are now universally acknowledged, with their tendency towards hallucination recognized as a fundamental reality. Galactica was, in essence, just a base model.
Here is a link to Ross Taylor's contribution to open source ( incidentally this app is created in part from the data @paperswithcode publishes - a startup he cofounded )
https://t.co/DYDgrggsic
I am the first author of the Galactica paper and have been quiet about it for a year. Maybe I will write a blog post talking about what actually happened, but if you want the TLDR:
1. Galactica was a base model trained on scientific literature and modalities.
2. We approached it with a number of hypotheses about data quality, reasoning, scientific modalities, LLM training, that hadn’t been covered in the literature - you can read about these in the paper.
3. For its time, it was a good model for its domain; outperforming PaLM and Chinchilla with 10x and 2x less compute.
4. We did this with a 8 person team which is an order of magnitude fewer people than other LLM teams at the time.
5. We were overstretched and lost situational awareness at launch by releasing demo of a *base model* without checks. We were aware of what potential criticisms would be, but we lost sight of the obvious in the workload we were under.
6. One of the considerations for a demo was we wanted to understand the distribution of scientific queries that people would use for LLMs (useful for instruction tuning and RLHF). Obviously this was a free goal we gave to journalists who instead queried it outside its domain. But yes we should have known better.
7. We had a “good faith” assumption that we’d share the base model, warts and all, with four disclaimers about hallucinations on the demo - so people could see what it could do (openness). Again, obviously this didn’t work.
8. A mistake on our part that didn’t help was people treated the site like a *product*. We put our vision etc on the site, which misled about expectations. We definitely did not view it as a product! It was a base model demo.
9. Pretty much every LLM researcher I’ve talked to (including at ICML recently) was complimentary about the strength of the research, which was sadly overshadowed by the demo drama - yes this was our fault for allowing this to happen.
10. Fortunately most of the lessons and work went into LLaMA 2; the RLHF research you see in that paper is from the Galactica team. Further research coming soon that should be interesting.
It’s a bit of a riddle because on the one hand the demo drama could have been avoided by us, but at the same time the “fake science” fears were very ridiculous and despite being on HuggingFace for a year, the model hasn’t caused any damage.
To reiterate: the anti-Galactica commentary was really stupid, however we should not have allowed that to even happen if we had launched it better.
I stick by the research completely - and even the demo decision, which was unprecedented openness for a big company with an LLM at the time, wasn’t inherently bad - but it was just misguided given the attack vectors it opened for us.
Despite all the above, I would do it all again in a heartbeat. Better to do something and regret, then not do anything at all. Still hurts though! 🙂