For example in convolutional neural networks (which are popular in computer vision) some architectures optimize over smooth loss surfaces and others do not.
A non smooth loss surface means that the loss values will bounce around and could spike + drop suddenly.
Finally for my ML peeps this is an oversimplified explanation. I know I didn’t delve into learning rates and LR scheduling (warm restarts in particular) and how they can impact loss values. Just trying to keep it high level for a wider audience.
Since this figure is going around, the sudden drop corresponds to a very standard learning rate decay and has nothing to do with MuonClip.
The point of this figure in the context of MuonClip was to show how stable the training was at such large scales
Today we’re announcing Meta LLM Compiler, a family of models built on Meta Code Llama with additional code optimization and compiler capabilities. These models can emulate the compiler, predict optimal passes for code size, and disassemble code. They can be fine-tuned for new optimizations and compiler tasks.
@HuggingFace repo ➡️ https://t.co/9URAr9sn5E
Research paper ➡️ https://t.co/nIYvWHqm1D
LLM Compiler achieves state-of-the-art results on code size optimization and disassembly. This work shows that AI is learning to optimize code and can assist compiler experts in identifying opportunities to apply optimizations.
We’re releasing LLM Compiler 7B & 13B models under a permissive license for both research and commercial use in the hopes of making it easier for developers and researchers alike to leverage this in their work and carry forward new research in this space.
guys going on podcasts talking about ai doom meanwhile they’re drawing straight lines on log charts… bro you’re worried about the wrong high school level intelligence
JUST LET ME SAY “TOKEN” ONE MORE TIME BRO
PLEASE JUST ONE MORE “TOKEN” BRO
I PROMISE THE INVESTORS WILL LOVE ME AGAIN BRO
I SWEAR BRO JUST ONE MORE, JUST LET ME SAY IT BRO
@teortaxesTex@nearcyan They’re sometimes better than traditional search and sometimes worse. But I think that the killer is that search is fast and optimized. It’s so cheap per query that Google can let you use it for free and ad support it.
Using an LLM for search is 1000s of times more inefficient.
@teortaxesTex@nearcyan Yeah I think entertainment is a big one people are using it for now. And for summarization, translation and other language manipulation too.
But tbh I’m kind of bearish on the search use case still. It’s not that LLMs aren’t good at it 1/2
I think you can apply that mindset to anything you do. And if you do, you’ll have the ability to capitalize on opportunities in a way that other people simply cannot.
TL;DR Prioritize depth & your own growth. Learn how to learn. Master yourself. And follow your passion. 7/7 fin.
It’s an overdone question and imo I think it’s actually the wrong one. No one else can tell you what to work on. You gotta find your personal product market fit! But also that’s dogshit advice because the next question is gonna be “how?” 🧵👇
I know this is an overdone question but, if you’re an experienced engineer/founder/investor what would you start working on today with your entire career ahead of you?
Jim Simmons has a great quote about how he hires people to work at RenTech that largely echoes this sentiment. “Just get people who have done real science, they don’t need to know finance.”
And to be clear this isn’t me telling you to go into academia or get a PhD. 6/n