I've consistently found that people who think they already have all the answers will always refuse to ingest new information that might contradict their views. They just don't want to risk their certitudes. As a result, they never learn anything new and new grow.
Atención #nlproc Latinoamérica:
Está abierta la convocatoria de @naacl de pequeños financiamientos para iniciativas regionales. Me ayudan a difundir por favor?
Deadline: 30 de abril
Más info: https://t.co/VQ7G4PzN70
Proyectos ya financiados:
https://t.co/c7w7jJSwrf
Our paper was accepted in @LrecColing !! Are you a Latinamerican doing research? Doesn't it feel like when your team marks a score in the Soccer World Cup or eliminatory? ⚽️ 🇵🇪🏆🥲
#LREC#signLanguageProcessing
It is only rarely that, after reading a research paper, I feel like giving the authors a standing ovation. But I felt that way after finishing Direct Preference Optimization (DPO) by @rm_rafailov@archit_sharma97@ericmitchellai@StefanoErmon@chrmanning and @chelseabfinn. This beautiful paper proposes a much simpler alternative to RLHF (reinforcement learning from human feedback) for aligning language models to human preferences.
RLHF has been a key technique for training LLMs. In brief, RLHF (i) Gets humans to specify their preferences by ranking LLM outputs, (ii) Trains a reward model (used to score LLM outputs) -- typically represented using a transformer network -- to be consistent with the human rankings, (iii) Uses reinforcement learning to tune an LLM, also represented as a transformer, to maximize rewards. This requires two transformer networks, and RLHF is also finicky to the choice of hyperparameters.
DPO simplifies the whole thing. Via clever mathematical insight, the authors show that given an LLM, there is a specific reward function for which that LLM is optimal. DPO then trains the LLM directly to make the reward function (that’s now implicitly defined by the LLM) consistent with the human rankings. So you no longer need to deal with a separately represented reward function, and you can train the LLM directly to optimize the same objective as RLHF.
Although it’s still too early to be sure, I am cautiously optimistic that DPO will have a huge impact on LLMs and beyond in the next few years.
You can read the paper here: https://t.co/m14qRYszVa I also write more about this in The Batch (linked to below).
https://t.co/8h2ag2plIa
Meet Mobile ALOHA! 🤖
What’s new:
- low-cost but widely-capable teleop platform 🦾🛞
- imitation learning just works
Paper/Code/Videos: https://t.co/I7Qm64cwK0
Led by @zipengfu and @tonyzzhao
Want to learn about meta-learning & few-shot learning?
All of the latest lecture videos for Stanford CS330 are now online!
https://t.co/PqNK2yqQ16
New topics in Fall '22 include:
- self-supervised pre-training
- large scale meta-optimization
- domain adaptation & generalization
Exciting times, welcome Gemini (and MMLU>90)! State-of-the-art on 30 out of 32 benchmarks across text, coding, audio, images, and video, with a single model 🤯
Co-leading Gemini has been my most exciting endeavor, fueled by a very ambitious goal. And that is just the beginning! A long 🐍 post about our Gemini journey & state of the field.
The biggest challenges in LLMs are far from trivial or obvious. Evaluation and data stand out to me. We've moved beyond the simpler "Have we won in Go/Chess/StarCraft?" to “Is this answer accurate and fair? Is this conversation good? Does this complex piece of text prove the theorem?” Exciting potential coupled with monumental challenges.
The field is less ripe further down the model pipeline. Pretraining is relatively well understood. Instruction tuning and RLHF, less so. In AlphaGo and AlphaStar we spent 5% of compute in pre-training and the rest in the very important RL phase, where the model learns from its successes or failures. In LLMs, we spend most of our time on pretraining. I believe there’s huge potential to be untapped. Cakes with lots of cherries, please 🎂
@Google has demonstrated its ability to move fast. It has been an absolute blast to see the energy from my colleagues and the support received. A “random” highlight is coauthoring our tech report with a co-founder. Another is coleading with @JeffDean. But beyond individuals, Gemini is about teamwork: it is important to recognize the collective effort behind such achievements. Picture a room full of brilliant people, and avoid attributing success solely to one person.
On a personal note, recently I celebrated my 10 year anniversary at Google, and it’s been 8 years since @quocleix and I co-authored “A Neural Conversational Model”, which gave us a glimpse of what was, has, and is yet to come. Back then, that line of work received a lot of skepticism. Lessons learned: whatever your passion is, push for it!
Zooming back out, there’s lots of change in our field, and the stakes couldn’t be higher. Excited for what’s to come from Gemini, but humbled by the responsibility to “get it right”. 2024 will be drastic. Welcome Gemini!
https://t.co/X4ZLpcUiiL
I'm recruiting PhD students for my lab at Johns Hopkins!
Please apply if you're interested in reliable ML / causal inference for decision-making in healthcare. See my website (https://t.co/vg0mn6gY5r) for more info.
Deadline 12/15. Retweets welcome :)
https://t.co/COLkwSjPNg
La Dra. Gissella Bejarano comparte con nosotros su apasionante labor en la creación del diccionario web de lengua de señas peruana @gissemari#Lenguadeseñas
I am recruiting two talented Ph.D. students to join me at
@KhouryCollege of Computer Sciences at
@Northeastern to work on cutting-edge NLP/HCI + Social Justice research. Deadline: Dec. 15th Application: https://t.co/tybrALccio
Interested in exploring fine-tuning? Join our Generative AI Study Group and discuss fine-tuning questions, share informative articles, trade fine-tuning techniques, and more! Meetup is every Friday at 8am PT. Feel free to invite a friend and register at https://t.co/TDhwjxmtZ6.
My greatest fear for the future of AI is if overhyped risks (such as human extinction) lets tech lobbyists get enacted stifling regulations that suppress open-source and crush innovation.
Read more in our Halloween special issue of the Batch: https://t.co/CNqP2TYBLM
We still see lots of links to old releases of CS224N. Make sure you're getting the latest goodness (RLHF, prompting, transformers) from the 2023 release!
YouTube: https://t.co/6hb0EIaR5Z
Website: https://t.co/5ReBrK9J4n
For fee cohort-based online class: https://t.co/bDh197g0X6
Durante mi estancia en la @pucp , he tenido la oportunidad de estar en contacto con investigadores del @iapucp , gracias al Prof. Dr. @CSARARMANDOBEL2 , con el cual vamos a colaborar en proyectos futuros. También participe en la impartición del Workshop https://t.co/oNbhOS4ekC
📢¡latin americans in AI!📢
pre-register before october 13th to attend a conference/summer-school in quito next february with some incredible speakers (pictured below).
don't think you can afford it? we have travel grants, so apply now!
https://t.co/WYS6cPTpIm
"The Cohere For AI Scholars Program is an 8-month, full-time research apprenticeship. The Scholars Program runs from January 8, 2024 - August 23, 2024. "
Apply 👇
🌟 Calling all bright minds from South America 🌎✨
With just 3% of applications coming from South America, we’re looking for more Scholars Program applicants in South America! If you’re in search of an opportunity to develop your research skills, your journey starts here.