instead of watching 2 hours of Netflix tonight, watch this Stanford lecture
it's the clearest explanation I've seen of how ChatGPT and Claude actually work
useful whether you've never touched AI in your life or have been using it every day for the past year
i took the key ideas and turned them into a practical guide on how to actually get 100% out of AI
you can find it below with ready-to-copy prompts and solutions
The buzz over DeepSeek this week crystallized, for many people, a few important trends that have been happening in plain sight: (i) China is catching up to the U.S. in generative AI, with implications for the AI supply chain. (ii) Open weight models are commoditizing the foundation-model layer, which creates opportunities for application builders. (iii) Scaling up isn’t the only path to AI progress. Despite the massive focus on and hype around processing power, algorithmic innovations are rapidly pushing down training costs.
About a week ago, DeepSeek, a company based in China, released DeepSeek-R1, a remarkable model whose performance on benchmarks is comparable to OpenAI’s o1. Further, it was released as an open weight model with a permissive MIT license. At Davos last week, I got a lot of questions about it from non-technical business leaders. And on Monday, the stock market saw a “DeepSeek selloff”: The share prices of Nvidia and a number of other U.S. tech companies plunged. (As of the time of writing, some have recovered somewhat.)
Here’s what I think DeepSeek has caused many people to realize:
China is catching up to the U.S. in generative AI. When ChatGPT was launched in November 2022, the U.S. was significantly ahead of China in generative AI. Impressions change slowly, and so even recently I heard friends in both the U.S. and China say they thought China was behind. But in reality, this gap has rapidly eroded over the past two years. With models from China such as Qwen (which my teams have used for months), Kimi, InternVL, and DeepSeek, China had clearly been closing the gap, and in areas such as video generation there were already moments where China seemed to be in the lead.
I’m thrilled that DeepSeek-R1 was released as an open weight model, with a technical report that shares many details. In contrast, a number of U.S. companies have pushed for regulation to stifle open source by hyping up hypothetical AI dangers such as human extinction. It is now clear that open source/open weight models are a key part of the AI supply chain: Many companies will use them. If the U.S. continues to stymie open source, China will come to dominate this part of the supply chain and many businesses will end up using models that reflect China’s values much more than America’s.
Open weight models are commoditizing the foundation-model layer. As I wrote previously, LLM token prices have been falling rapidly, and open weights have contributed to this trend and given developers more choice. OpenAI’s o1 costs $60 per million output tokens; DeepSeek R1 costs $2.19. This nearly 30x difference brought the trend of falling prices to the attention of many people.
The business of training foundation models and selling API access is tough. Many companies in this area are still looking for a path to recouping the massive cost of model training. Sequoia’s article “AI’s $600B Question” lays out the challenge well (but, to be clear, I think the foundation model companies are doing great work, and I hope they succeed). In contrast, building applications on top of foundation models presents many great business opportunities. Now that others have spent billions training such models, you can access these models for mere dollars to build customer service chatbots, email summarizers, AI doctors, legal document assistants, and much more.
Scaling up isn’t the only path to AI progress. There’s been a lot of hype around scaling up models as a way to drive progress. To be fair, I was an early proponent of scaling up models. A number of companies raised billions of dollars by generating buzz around the narrative that, with more capital, they could (i) scale up and (ii) predictably drive improvements. Consequently, there has been a huge focus on scaling up, as opposed to a more nuanced view that gives due attention to the many different ways we can make progress. Driven in part by the U.S. AI chip embargo, the DeepSeek team had to innovate on many optimizations to run on less-capable H800 GPUs rather than H100s, leading ultimately to a model trained (omitting research costs) for under $6M of compute.
It remains to be seen if this will actually reduce demand for compute. Sometimes making each unit of a good cheaper can result in more dollars in total going to buy that good. I think the demand for intelligence and compute has practically no ceiling over the long term, so I remain bullish that humanity will use more intelligence even as it gets cheaper.
I saw many different interpretations of DeepSeek’s progress here in X, as if it was a Rorschach test that allowed many people to project their own meaning onto it. I think DeepSeek-R1 has geopolitical implications that are yet to be worked out. And it’s also great for AI application builders. My team has already been brainstorming ideas that are newly possible only because we have easy access to an open advanced reasoning model. This continues to be a great time to build!
[Original text: https://t.co/yiOHeGJgLZ ]
Yesterday, my Twitter finally got 25,000+ followers🥳
I've been preparing for this day for several months now, and today I present to you the book
PYTHON FOR OSINT. 21 DAY COURSE FOR BEGINNERS
https://t.co/9ntDwKrRi5
(the course is free, donate just if you want)
Let’s go Oilers! 🏒 #EPSB students and staff have been sporting their Oilers gear this playoff run and cheering on the home team! Check out these photos from across the Division! #yeg#LetsGoOilers
The wait is over 🥳 I completed some major updates for VMwareCloak, a Powershell script that sanitizes VMware Workstation VM's from some of the common VM-detection techniques used by #malware! Try it out and let me know how it works for you 👇
https://t.co/yStH40sDs9
Is malware detecting your VirtualBox VM's? Is pafish giving you trouble? Try out the latest release of my PowerShell-based tool VBoxCloak! A quick and dirty way to hide your VM's from some common VM-detection techniques. VMware coming soon! 🥳
https://t.co/dRrnp7O2kw
AvatarAPI
Enter email address and receive an image of the avatar linked to it.
Over a billion avatars in the database collected from public sources (such as Gravatar, Stackoverflow etc.)
https://t.co/66e4XIBRZV
n0kovo_subdomains
Wordlist for subdomain enumeration of 3,000,000 lines, crafted by harvesting SSL certificates from the entire IPv4 space. Shortened versions of the list are also available: 1 000 000, 500 000, 200000 and 50000 lines
https://t.co/PavaswPx4v
Contributor @n0kovo
Recently advancements in AI/ML technology are changing our world. To keep up with the disruption, we have been working on a tool to solve complex problems with ATT&CK.
Our attackgpt Twitter bot is ready for beta testing. You can interact with it by using the hashtag #attackgpt.
8 pieces of free software for cybersecurity enthusiasts:
1. Training: Hack The Box
2. Curated News: Feedly
3. Web Hacking: Burp Suite
4. Data Modification: Cyber Chef
5. Port Scan: Nmap
6. Operating System: Kali Linux
7. Debugging: Ghidra
8. Email Security: DeHashed
Alpaca LoRA Playground
You've probably heard about the "ChatGPT copy" created by Stanford University researchers, which cost less than US$600 to train up and #opensource (https://t.co/z2xEnNPQvb)
Here you can try it in action:
https://t.co/3LjeniozfD
Glad to announce that we will be back (virtually) with two classes at @BlackHatEvents ⚡️
Azure Attacks for Red and Blue Teams
and
Active Directory Attacks for Red and Blue Teams
#redteam#BHUSA#Azure
https://t.co/uv7bdy2SQH
https://t.co/mnjYGxVwoT
ChatGPT is changing the game, and I want to share real things you can do with this AI system today.
Please save this thread and start testing this technology NOW so you’re ahead of the curve.