So if you are typical ML researcher, you had this question for eternity:
"I want small, powerful model: Should we train large model and distill? Or should we train small model from scatch"
This new Apple papers conclusion:
Its complicated but maybe yes, depending on your budget.
1/n
New research shows that LLMs don't perform well on long context.
Perfect needle-in-the-haystack scores are easy—attention mechanisms can match the word. When you require 1-hop of reasoning, performance degrades quickly.
This is why guaranteeing correctness for agents is hard.
Google may be only a year or two away from total disruption. AI will eliminate the Search Engine Result Page, which is where they make most of their money.
Even if they catch up on AI, they can't fully deploy it without destroying the most valuable part of their business!
What would the worker need in order to escape the hole? Centripetal acceleration = v^2/r: increasing r means you have to increase v to compensate, so more velocity and a larger frictional force [read more: https://t.co/7TORXuY9rL]
"C'est mort. Pour moi, c'est mort depuis la sortie du troisième rapport du GIEC le 4 avril dernier." @gemenne au sujet de la question écologique.
La suite :
➡️ https://t.co/KtYgK3iUJD
🎧 en podcast
I never ask this, but I'm asking now: Please retweet this.
Texas plans to kill Melissa Lucio in four days. Five jurors say evidence was withheld from them. A bipartisan majority of the Texas legislature favors clemency. Why are you waiting, @GovAbbott?
https://t.co/9iMkVhNDrQ
These are real intercepted calls: Russian soldiers in Ukraine call their close ones back in Russia to tell how it is going so far. Looting and war crimes included. Please, share! The world must know the truth of what they’re doing to our homes and people.
@paulg quote you will enjoy from our new python developer intern @issabanzan : "I like to code, it is like playing a game. When you're stuck at some point, you look for a solution to move forward everywhere in the game and on the forums"