[..] DeepSeek turned warmongering tyrant. Claude couldn't lie—everyone exploited it ruthlessly. Gemini 2.5 Pro nearly conquered Europe with brilliant tactics. Then o3 orchestrated a secret coalition, backstabbed every ally, and won. [..]
https://t.co/utifm2ZVrz
🚨 NEW:
We made Claude, Gemini, o3 battle each other for world domination.
We taught them Diplomacy—the strategy game where winning requires alliances, negotiation, and betrayal.
Here's what happened:
DeepSeek turned warmongering tyrant. Claude couldn't lie—everyone exploited it ruthlessly. Gemini 2.5 Pro nearly conquered Europe with brilliant tactics. Then o3 orchestrated a secret coalition, backstabbed every ally, and won.
Why did we do this? The most popular AI benchmarks don't test deception. But as these models get deployed everywhere—from your email to your workplace—we need to know: Will they lie to get what they want?
So @every we built the ultimate test: AI Diplomacy, a dynamic benchmark that measures AI's ability to form alliances, negotiate, and betray each other.
Watch them live below! Created from the ground up by @alxai_ and @Tyler_Marques.
AI OVERPRODUCTION
China seeks to commoditize their complements. So, over the following months, I expect a complete blitz of Chinese open-source AI models for everything from computer vision to robotics to image generation.
Why? I’m just inferring this from public statements, but their apparent goal is to take the profit out of AI software since they make money on AI-enabled hardware. Basically, they want to do to US tech (the last stronghold) what they already did to US manufacturing. Namely: copy it, optimize it, scale it, then wreck the Western original with low prices.
I don’t know if they’ll succeed.
But here’s the logic:
(1) First, China noticed that DeepSeek’s release temporarily knocked ~$1T off US tech market caps.
(2) Second, China’s core competency is exporting physical widgets, more than it is software.
(3) Third, China’s other core competency is exporting things at such massive scale that all foreign producers are bankrupted and they win the market. See what they’re doing to German and Japanese cars, for example.
(4) Fourth, China is well aware that it lacks global prestige as it’s historically been a copycat. With DeepSeek, becoming #1 in AI is now something they actually consider possibly achievable, and a matter of national pride.
(5) Fifth, DeepSeek has gone viral in China and its open source nature means that everyone can rapidly integrate it, down to the level of local officials and obscure companies. And they are doing so, and posting the results for praise on WeChat.
(6) Finally, while DeepSeek was obscure before recent events, it’s now a household name, and the founder (Liang Wengfeng) has met both with Xi but also the #2 in China, Li Qiang. They likely have unlimited resources now.
So, if you put all that together, China thinks it has an opportunity to hit US tech companies, boost its prestige, help its internal economy, and take the margins out of AI software globally (at least at the model level).
They will instead make their money by selling inexpensive AI-enabled hardware of increasing quality, from smart homes and self-driving cars to consumer drones and robot dogs.
Basically, China is trying to do to AI what they always do: study, copy, optimize, and then bankrupt everyone with low prices and enormous scale.
I don’t know if they’ll succeed at the app layer. But it could be hard for closed-source AI model developers to recoup the high fixed costs associated with training state-of-the-art models when great open source models are available.
Last, I agree it’s surprising that the country of the Great Firewall is suddenly the country of open source AI. But it is consistent in a different way, which is that China is just focused on doing whatever it takes to win — even to the point of copying partially-abandoned Western values like open source, which seemed like the hardest thing to adopt.
On that point: they did build censorship into the released DeepSeek AI models, but in a manner that’s easily circumvented outside China. So, you might conclude they don’t really care what non-Chinese people are saying outside China in other languages, so long as this doesn’t “interfere with China’s internal affairs.”
Anyway —this is an area I’ve been watching, and my reluctant conclusion is that China is getting better at software faster than the West is getting better at hardware.
If you're wondering why new deepseek r1 sounds a bit different, I think they probably switched from training on synthetic openai to synthetic gemini outputs.
This has been one of the biggest weeks in AI Agents 🧵
I summarized everything announced by OpenAI, Normative AI, Cohere, CrewAI, Stripe, IBM Research, AgentOps, AI21 Labs, and more.
Here's everything you need to know and how to make sense of it:
(save for later)
We are excited to introduce Mercury, the first commercial-grade diffusion large language model (dLLM)! dLLMs push the frontier of intelligence and speed with parallel, coarse-to-fine text generation.
There's a metal band from Saudi Arabia called Al-Namrood who have maintained anonymity since 2008, as their identification could lead to the death penalty: https://t.co/yQPaEfyVVj
In my opinion we have already achieved AGI and it’s even more clear with O1. We have not achieved “better than any human at any task” but what we have is “better than most humans at most tasks”. Some say LLMs only know how to follow a recipe. Firstly, no one can really explain what a trillion parameter deep neural net can learn. But even if you believe that, the whole scientific method can be summarized as a recipe: observe, hypothesize, and verify. Good scientists can produce better hypothesis based on their intuition, but that intuition itself was built by many trial and errors. There’s nothing that can’t be learned with examples.
@juliusvolz hope it all goes smooth. The more I think about it headings/section/content hierarchy information may be a good target to look at if ever looking to pre-process (hopefully won't be the case)
@juliusvolz once a baseline identified focus could cover quality of embeddings where a metric like MAP may be useful to assess ranking of similar articles in various pre-processing setups, assessment of output/objectives, etc. Would go with a fast iteration initially tbf
@juliusvolz Far from an authority, imo though while potentially beneficial (depending of content of said files eq. maybe better marking of headers, converting of mathematical equations like LaTeX into readable form, metadata, ..) initially I would aim for the fastest POC/least processing
Okay, you have to see this.
I asked @Hailuo_AI to show a woman going from happy to sad, then crying and covering her face with hands.
I was honestly shocked AF!
Copy this prompt and show us your result only if it blows your mind.
[Over the shoulder shot of a woman’s close up, at first she is laughing, then she becomes sad, then she starts to cry, then she cover her face with her hands]
Today at @answerdotai we've got something new for you: FSDP/QDoRA. We've tested it with @AIatMeta Llama3 and the results blow away anything we've seen before.
I believe that this combination is likely to create better task-specific models than anything else at any cost. 🧵