News: @Alibaba_Qwen Qwen-Max jumps to #7, surpassing DeepSeek-v3! 🔥
Highlights:
- Matches top proprietary models (GPT-4o/Sonnet 3.5)
- +30 pts vs DeepSeek-v3 in coding, math, and hard prompts
@ChatGLM GLM-4-Plus also breaks into top-10, Chinese AI companies are closing the gap fast! More analysis👇
Nunca imaginei que fosse ver um time treinado por Pep Guardiola ser tão frágil, exposto, fraco defensivamente. Tudo espaçado. Sem proteção. Frágil. Já tomou goleadas piores, mas ser dominado dessa forma pelo time do seu aprendiz é um golpe muito duro…
Finally took time to go over Dario's essay on DeepSeek and export control and to be honest it was quite painful to read. And I say this as a great admirer of Anthropic and big user of Claude*
The first half of the essay reads like a lengthy attempt to justify that closed-source models are still significantly ahead of DeepSeek. However, it mostly refers to internal unpublished evals which limit the credit you can give it, and statements like « DeepSeek-V3 is close to SOTA models and stronger on some very narrow tasks » transforming in a general conclusion « DeepSeek-V3 is actually worse than those US frontier models — let’s say by ~2x on the scaling curve » left me generally doubtful. The same applies to the takeaway that all discoveries and efficiency improvements of DeepSeek have been discovered long ago by closed-models companies, this statement mostly resulting from a comparison of DeepSeek openly published $6M training numbers with some vague « few $10M » on Anthropic side without providing much more details. I have no doubts the Anthropic team is extremely talented and I’ve regularly shared how impressed I am with Sonnet 3.5 but this longwinded comparison of open research with vague closed research and undisclosed evals has left me less convinced of their lead than I was before I reading it.
Even more frustrating was the second half of the essay which dive into the US-China race scenario and totally misses the point that the DeepSeek model is open-weights, and largely open-knowledge due to its detailed tech report (and feel free to follow Hugging Face’s open-r1 reproduction project for the remaining non-public part: the synthetic dataset). If both DeepSeek and Anthropic models had been closed source, yes the arm-race interpretation could have make sense but having one of the model freely widely available for download and with detailed scientific report renders the whole « close-source arm-race competition » argument artificial and unconvincing in my opinion.
Here is the thing: open-source knows no border. Both in its usage and its creation.
Every company in the world, be it in Europe, Africa, South-America or the USA can now directly download and use DeepSeek without sending data to a specific country (China for instance) or depending on a specific company or server for running the core part of its technology.
And just like most open-source library in the world are typically built by contributors from all over the world, we’ve already seen several hundred derivative models on the Hugging Face hub created everywhere in the world by teams adapting the original model to their specific use cases and explorations.
What's more, with the open-r1 reproduction and the DeepSeek paper, the coming months will clearly see many open-source reasoning models being released by teams from all over the world. Just today, two other teams, AllenAI in Seattle and Mistral in Paris both independently released open-source base models (Tülu and Small3) which are already challenging the new state-of-the-art (with AllenAI indicating that its Tülu model surpasses the performance of DeepSeek-V3).
And the scope is even much broader than this geographical aspect. Here is the thing we don’t talk nearly enough about: open-source will be more and more essential for our… safety!
As AI becomes central to our lives, resiliency will increasingly become a very important element of this technology. Today we’re dependent on internet access for almost everything. Without access to the internet, we lose all our social media/news feeds, can’t order a taxi, book a restaurant, or reach someone on WhatsApp. Now imagine an alternate world to ours where all the data transiting through the internet would have to go through a single company’s data centers. The day this company suffers a single outage, the whole world would basically stop spinning (picture the recent CrowdStrike outage magnified a millionfold).
Soon, as AI assistants and AI technology permeate our whole life to simplify many of our online and offline tasks, we (and companies using AI) will start to depend more on more on this technology for our daily activities and we will similarly start to find annoying or even painful any downtime in these AI assistants from outages.
The most optimal way to avoid future downtime situations will be to build resilience deep in our technological chain.
Open-source has many advantages like shared training costs, tunability, control, ownership, privacy but one of its most fundamental virtue in the long term –as AI becomes deeply embedded in our world– will likely be its strong resilience. It is one of the most straightforward and cost-effective ways to easily distribute compute across many independent providers and to even run models locally and on device with minimal complexity.
More than national prides and competitions, I think it’s time to start thinking globally about the challenges and social changes that AI will bring everywhere in the world. And open-source technology is likely our most important asset for safely transitioning to a resilient digital future where AI is integrated into all aspects of society.
*Claude is my default LLM for complex coding. I also love its character with hesitations and pondering, like a prelude to the chain-of-thoughts of more recent reasoning models like DeepSeek generations.
Our latest open 1M model is here! Context length is a fundamental feature of the LLM, and we've also provided some detailed technical insights. Enjoy it!
Interesting: whenever I read an "I've replaced N SaaS services with AI, in a day, saving $$$ per year, SaaS is dead" post, both are true:
1. The SaaS "replaced" is not named
2. The "AI replacing it" is not shown (no code, no specifics, no nothing)
Be careful what you believe
🇵🇹 Jorge Jesus falou abertamente sobre as condições físicas de Neymar após a vitória do Al-Hilal no Campeonato Saudita:
“O que é verdade, é que fisicamente ele não tem conseguido acompanhar a equipe”.
🎥: @cahemota
A POLÍCIA E O PCC
Nos últimos cinco anos, 111 policiais oriundos das polícias das cinco regiões do Brasil se tornaram réus ou foram condenados por envolvimento com a maior facção do país, revela levantamento inédito que eu e @alineamribeiro publicamos hoje no @JornalOGlobo.
🎯 Quem controla a distribuição controla o acesso ao público, independentemente da qualidade do conteúdo.
💡Permite também controlar como o conteúdo é apresentado e consumido, influenciando a percepção do usuário e a fidelização.
👇🏾”A principal linha de investigação é a de que Antônio Martins Santos Filho teria organizado uma chacina depois que lideranças do movimento se recusaram a permitir a tomada de um lote do assentamento negociado por criminosos.”
https://t.co/olahRwhsu1
"Não se trata de terra invadida, propriedade privada invadida. Pelo contrário: é um território regularizado pelo Incra e, portanto, invadido e atacado por criminosos. Esses sim, criminosos", comenta @flaviaol sobre ataque a assentamento do MST.
➡ Assista ao #Edição18: https://t.co/mgnXyoxnmY #GloboNews
📢 Após uma live de Elon Musk com o líder da ultradireita alemã, dezenas de universidades da Alemanha e Áustria anunciaram que estão abandonando o X.
"Os acontecimentos no X mostram que a plataforma não cumpre mais sua responsabilidade de promover um discurso justo. Como instituições acadêmicas, não podemos aceitar isso", disse a reitora da Universidade de Düsseldorf.
vai acabar, tudo isso vai acabar!
tu acha bonito, eu sei, mas não basta - seu like aqui não paga boleto de ninguém daí.
o senso de urgência é real - prestigie, gaste seu dinheiro, chame os amigos, saia do fetiche e viva o mundo real.
(+)
Here's the "Elon PoE2 debacle" story in written form on PCgamer if you prefer to read (and miss the hilarious facial expressions): https://t.co/UX7LBGnm9W
What an ego-trip... Elon Musk has been livestreaming himself playing overpowered Path Of Exile 2 character, proudly and seriously pretending that he got there _himself_, and actual POE veterans are laughing their asses off watching and dissecting that charade. Why Elon, why?🤦♂️🤦♂️