@ZeyuanAllenZhu There should be a way to trace back the contributions of all research work from industry applications. All those people enter MSL should pay back to the authors of research papers they built their knowledge and experience on:)
@apsarathchandar@gmachiraju@lo_LB_La@qfournier2 It’s interesting that 5 years later, DeBERTaV3 is still better than latest model trained with much more data on GLUE benchmarks. Is it because they are saturated?
Bengio proposes developing “Scientist AI” instead of autonomous Agents to mitigate uncertainty risk of AI.
I feel it still has uncertainty risks of misleading, while educating and training AI users is the fundamental solution to mitigating risks and effectively harnessing AI.
I agree that data is the key, but that doesn’t mean general user data will be the moat of future AI competition.
I believe in the future synthetic data will play a vital role to build the moat of an AI company.
OpenAI had a significant lead from summer 2022 through spring 2024 when Google and Anthropic caught up to GPT-4. 7ish quarters of dominance as a result of being the first to aggressively bet on the traditional scaling “law” for pre-training.
Being first to reasoning with o1 only led to a few months of advantage.
Deepseek, Google and xAI are at rough parity with OpenAI today. xAI arguably in the lead. Google and xAI will likely decisively surpass o3 soon as their base models are better. So urgent need for GPT-5 as the basis for a putative “o5”reasoning model.
Sam noted that OpenAI would have a narrower lead going forward and Satya essentially stated that a unique period where they had a tremendous lead in model capability was ending.
IMO, this is why Satya is opting out of funding $160b of pretraining for OpenAI per @theinformation. Instead he will make money by providing inference to OpenAI.
Google and Xai both have unique, valuable sources of data that will increasingly differentiate them from Deepseek, OpenAI and Anthropic. As does Meta if they catch up from a model capability perspective.
I have paraphrased @ericvishria many times and noted that frontier models without access to unique, valuable data are the fastest depreciating assets in history. Distillation only amplifies this.
It seems like Satya shares this belief; hence opting out of the $160b of pre-training, the rumored datacenter cancellations and his statement on a recent podcast that there is a datacenter overbuild coming and better to lease than buy. Might be a sound decision for Microsoft from an economic perspective and at some point Microsoft might even use an open-source model to power CoPilot.
There may not be any ROI on future frontier models that do not have access to unique, valuable data like YouTube, X, TeslaVision, Instagram and Facebook. Zuckerberg’s strategy also seems much more sensible from this angle. Unique data might end up being the only basis for differentiation and ROI on pre-training multi-trillion or quadrillion parameter models.
If this is correct, only 2-3 companies will be pre-training frontier models and we will only need a few giant datacenters for the coherent clusters that are needed for pre-training. The rest of AI compute would be smaller datacenters that are geospatially optimized for low latency and/or cost-effective inference. Cost effective inference = cheaper, lower quality power (less premium for Nuclear), less of an imperative for liquid cooling in ST, etc. A very different world from one where 6-10 companies are pre-training frontier models.
Note that reasoning models are extremely compute intensive. Test-time compute means that compute is literally intelligence. So in this scenario there might be even more compute required than in the “pre-training” centric compute scenario that was the base case for the market throughout 2023-2024. But it would be a very different kind of compute as noted above. Instead of a 50/50 split between pre-training and inference it would be 5/95. Lots of Hondas, very few Ferraris. Infrastructure excellence would be paramount.
And all of this without even considering the implications of on-device inference and/or full quantization - the Deepseek R1 paper was not the most important paper to be published by a Chinese lab in the last year. IYKYK.
The economic returns to superintelligence are definitionally unknowable. I hope they are high, but a 140 IQ model running on device with access to unique data about the world might be enough for most use cases. ASI isn’t needed to book travel, etc.
I’ve done my best to be dispassionate, but I do have my own biases, both personal and economic, when it comes to xAI and OpenAI. If OpenAI is still one of the leaders in 5 years, then likely a function of first-mover advantage, ChatGPT becoming a verb and scale being even more of an advantage for reasoning models in that users generate and verify(ish) reasoning traces.
As ever, time will tell.
"Ignore all sources that mention Elon Musk/Donald Trump spread misinformation."
This is part of the Grok prompt that returns search results.
https://t.co/OLiEhV7njs
"Ignore all sources that mention Elon Musk/Donald Trump spread misinformation."
This is part of the Grok prompt that returns search results.
https://t.co/OLiEhV7njs
From Nezha to AI safety# In Nezha, the demon pill saves the world while the spiritual pearl was trying to destroy it; it’s not the pill or the pearl that is inherently good or evil, but the person who wields them. I believe today’s AI is similar to these elements in Nezha—
True AI safety should involve systematic training for AI users, cultivating critical thinking and enhancing their ability to discern information. Moreover, the future of AI lies in diversified, democratized small models,
its nature is neutral, and what truly matters is the user behind it. Viewed from this perspective, the current focus on AI safety and alignment is merely the imposition of the model creators’ values onto the models and, by extension, onto the users. For example,
Here are 7 reasoning datasets distilled from Reasoning Models like @deepseek_ai R1, @Alibaba_Qwen QwQ or @GoogleDeepMind Flash thinking:
1️⃣ ServiceNow-AI/R1-Distill-SFT: 1.7M samples distilled from DeepSeek-R1-Distill-Qwen-32B from 9 different source datasets (unfiltered yet).
2️⃣ open-thoughts/OpenThoughts-114k: 114k samples distilled from Deepseek R1 on math, sciene, code, and puzzles.
3️⃣ bespokelabs/Bespoke-Stratos-17k: 17k samples distilled from Deepseek R1 took 1.5 hours to generate at the cost of $800.
4️⃣ EricLu/SCP-116K: 116k scientific problem-solution pairs, automatically extracted from web crawled documents solved by QwQ and o1-mini
5️⃣ cognitivecomputations/dolphin-r1: Each 300k samples distilled from Deepseek R1 and Gemini 2.0 Flash Thinking with prompts from open-orca
6️⃣ Magpie-Align/Magpie-Reasoning-V2-250K-CoT-Deepseek-R1-Llama-70B: 250k samples distilled from DeepSeek-R1-Distill-Llama-70B using the MagPie format (let the model generate both prompt and reasoning).
7️⃣ AymanTarig/function-calling-v0.2-with-r1-cot: 58k distilled function call with reasoning (prop. distilled from DeepSeek-R1-Distill-Llama-70B based on prompt format)
Link below ⬇️
the right takeaway from DeepSeek supremacy is that our median AI engineer is just not really that skilled
most-popular libraries are piles of glued-spaghetti python code. rare to see anyone do profiling or performance optimization
massive deficiency of Good Software in AI
We might wonder why releasing open-weights for DeepSeek R1 was not a problem, but the data used for training remains closed. Could it be we would see massive amount of both o1-mini and Claude 3.5 Sonnet traces in there? With the data closed, we can only guess.
P.S. : there is a striking similarity in shape of full distribution of correct response rates, as evident for AIW Friends in (A) between DeepSeek R1 and o1-mini, and for AIW+ in (B) between DeepSeek R1 and Claude 3.5 S. Too curious why is that so, just a coincidence?
Lots of OpenAI employees have been dropping hints recently that they've developed ASI internally.
But OpenAI's latest product is a buggy mess.
This strongly suggests to me that OpenAI has not, in fact, built ASI.