❓Can I collect some feedbacks: Is fully open-source research necessary?
Earlier, I released a family of 1-8B models (open data, code, weights): beating Llama3-8B with <10% pretrain time, beating most (all?) open-data models of this scale.
🔓No shortcuts: 10+ legal debates for using open data, 10+ more for weights, months of blockers, endless nights scavenging GPUs.
📈All to provide a strong, reproducible baseline — one I believed critical for understanding the physics of LLMs. Next planned was GLA + Canon, outperforming all modern linear models in tests.
Yet attention was low.
❓Should I close-source to save time and focus on pure research? Honest feedback appreciated.
Test of Time Winner
Adam: A Method for Stochastic Optimization
Diederik P. Kingma, Jimmy Ba
Adam revolutionized neural network training, enabling significantly faster convergence and more stable training across a wide variety of architectures and tasks.
@iclr_conf many many many thanks to @kchonyc and @Yoshua_Bengio for enabling the wildest ever start of my research career
2014 was a very special time to do deep learning, a commit that changes 50 lines of code could give you a ToT award 10 years later 😲
@srush_nlp Same here. Email is also slow. Download a kindle app on iPhone, open the arxiv link pdf file version on iPhone, share to the kindle app through Safari sharing button, almost immediately appeared on Kindle Scribe. Only issue to improve is that scribe size larger would be better
@Wenxuan_Zhou I think your suggestion is a better approach based on the conclusions in recent amazing research from @ZeyuanAllenZhu Physics of Language Models
🏆Exciting that our Mixture of Agents (MoA) tops the AlpacaEval leaderboard!
We introduce the MoA architecture: layers of diverse LLM agents fuse + improve prev LLMs' outputs. MoA of only open-source LLMs outperforms GPT-4o by 7% on AlpacaEval2 while more cost-efficient🚀1/5
How can we train full-size humanoid robots?
New paper introducing:
- learned controller for shadowing humans
- imitation learning of demos collected via shadowing
Website with code & videos: https://t.co/uX2aEahPCL
We're sharing progress toward understanding the neural activity of language models. We improved methods for training sparse autoencoders at scale, disentangling GPT-4’s internal representations into 16 million features—which often appear to correspond to understandable concepts.
https://t.co/tTRPztmra1
MLSys is around the corner. Catch @Azaliamirh talk at the #MLSys2024 Young Professionals Symposium! May 13th in Santa Clara, California. Full agenda here: https://t.co/ZVo1TVf3a6
Catch @KurtKeutzer keynote talk at the #MLSys2024 Young Professionals Symposium! May 13th in Santa Clara, California. Full agenda here: https://t.co/9oNuFRTJai
It's been a wild ride. Just 20 of us, burning through thousands of H100s over the past months, we're glad to finally share this with the world! 💪
One of the goals we’ve had when starting Reka was to build cool innovative models at the frontier. Reaching GPT-4/Opus level was a personal goal for many of us in the team. Doing it from scratch, on top of starting a company, makes it even more challenging but rewarding. 😁
Core is still improving (not done training!) but we’re happy to ship an early version 🚢. I’ve been vibe-checking it for a bit and it’s a really cool model (especially at multimodal) 😎.
Check out the blogpost, technical report and very non-cherry picked, “in the wild” showcase/demo in the thread below! Core is competitive with true frontier models. It beats Claude3 Opus on multimodal chat and matches GPT4-V on MMMU. Text metrics are competitive too (~83+ MMLU). In my mind, this is our arrival at the frontier. 😎👌🔥
More fun stuff to come in the following weeks! 😋
Happy International Women's Day!
We're celebrating women who have changed the world. Here's all of the amazing women who have received the #NobelPrize and their remarkable achievements at the time of the award.
Who are the women who inspire you the most?
#IWD2024