QLoRA: 4-bit finetuning of LLMs is here! With it comes Guanaco, a chatbot on a single GPU, achieving 99% ChatGPT performance on the Vicuna benchmark:
Paper: https://t.co/J3Xy195kDD
Code+Demo: https://t.co/SP2FsdXAn5
Samples: https://t.co/q2Nd9cxSrt
Colab: https://t.co/Q49m0IlJHD
@nucc@ELLISforEurope In some countries, yes. But some countries do allow pursuing a PhD with a BSc. Or at least this is what they say, it looks they don't actually mean it.
@mido_assran This paper to semi-supervised learning is a treasure. Not using labels directly for supervised training, but using them to provide a nice uniform distribution over the predictions instead. Very nice work! I wonder how far representation learning will go in the next few years.
First day @ #MIPS2019 for Asura is a hit! Time and time again it is our live demos of vehicle recognition using make and model recognition and #LPR all integrated into #Milestone XProtect drawing lots of interest and opportunities for new acquaintances. Getting ready for day 2!
Such a multifunctional course from https://t.co/rtRK421Twv!!! If you ever wanted to be the master of deep learning and the rockstar of drawing, @AndrewYNg is your guy! 😉 #DeepLearning#NeuralNetworks
Been experimenting with PPO+Curiosity on a more sparse-rewarding environment. Agent has to learn to press the button to spawn a pyramid, then knock the pyramid over to get the (+1 reward) gold brick. Vanilla PPO completely fails. PPO+Curiosity solves it. 🎉
Kickstarting Deep Reinforcement Learning proposes a paradigm where 'teacher' agents help train 'student' agents. Benefits include faster research cycles and students that can surpass their teachers: https://t.co/QwWd3kxV7B
goodbye stephen hawking now that the stars can finally meet you i think that they will be so proud that their atoms created someone quite as special and as brilliant as you xox
Do you want to feel as dumb as an RL agent? Try playing this game: https://t.co/ukROIkFuKn where we tried to remove many of the priors that humans typically use. Checkout our paper alongwith @deepakpathak at https://t.co/90SrkYq3Dm
Spread the word fam
This is a remote program and is open to anyone with US work authorization located in US timezones (we’re happy to provide a desk in our San Francisco office if you happen to be located in the Bay Area).
https://t.co/EAP0ziiRu1