🥁 Llama3 is out 🥁
8B and 70B models available today.
8k context length.
Trained with 15 trillion tokens on a custom-built 24k GPU cluster.
Great performance on various benchmarks, with Llam3-8B doing better than Llama2-70B in some cases.
More versions are coming over the next few months.
https://t.co/EkU9aIHdZE
Introducing Gemini 1.5: our next-generation model with dramatically enhanced performance. It also achieves a breakthrough in long-context understanding.
The first release is 1.5 Pro, capable of processing up to 1 million tokens of information. 🧵 https://t.co/qT0aXdFL0n
@gdb Interesting that the president of OpenAI still can spend all day coding. It’s important that staff in all levels of software and data science leadership rotate in and out of development or research work. Leaders too removed from hands on work become out of touch and ineffective.
🤩 Lowkey Goated When #Classifier Is The Vibe! 🔥 Check out this amazing paper by Wenyang Liu et al. including @wangyi8848 to learn more about CNN for File Fragment Classification Using Bit Shift and n-Gram Embeddings https://t.co/K1lehrZZni
Unveiling GPT-4 -- our large multimodal model that exhibits human-level performance on various professional and academic benchmarks. With iterative alignment and adversarial testing, it's our best-ever model on factuality, steerability, and safety.
https://t.co/rjsIYWTN3Y
Lowkey Goated When Counting Objects Without Labels is the Vibe! Check out this groundbreaking paper by @nguyentienvu, Jingyi Xu et al. #AI#ObjectDetection#ZeroShotLearning 🤯🤩 Link: https://t.co/Y1J0IHOL9p
@ericweinstein You know he bought Twitter for the political power it gives him. Next step was to get rid of anyone with a backbone, who may oppose any kind of morally dubious manipulation of the public. Then control of the narrative can begin. Why assume the poll was even real..?
Can language model pretraining be even better?
Our paper shows that by randomly masking input tokens during pretraining, the zero-shot, few-shot, and fine-tuning performance can be significantly improved.
https://t.co/Jw2gTMLKPK
🧵
Have you ever wondered whether higher education can be an important driver of growth in developing countries?
My job market paper tries to answer this question by looking at a national expansion of higher ed in Vietnam.
Over 100 new universities were opened during 2006-2013!
How much data are augmentations worth? We show that augmentations can actually be worth more than extra data and invariance! They increase variance across batches, and this extra stochasticity finds flatter minima. https://t.co/wn1pL7RNbY 1/8
Introducing UL2, a novel language pre-training paradigm that improves performance of language models across datasets and setups by using a mixture of training objectives, each with different configurations. Read more and grab model checkpoints at https://t.co/A7ZAFNMCY6