5 years ago i was giving an interview where the guy grilled me for not answering a textbook question word-for-word. anyway, later on the call he accidentally opened Vim and literally had to kill the whole terminal because he couldnโt figure out how to exit ๐ญ
5 years ago i was giving an interview where the guy grilled me for not answering a textbook question word-for-word. anyway, later on the call he accidentally opened Vim and literally had to kill the whole terminal because he couldnโt figure out how to exit ๐ญ
DeepSeek just completely scrambled the AI leaderboards!
Based on the benchmarks and the latest API updates, here is the massive leap they just made:
- Crushes GPT-5.5: Scores 70.3 vs 55.6 on Toolathlon for long-horizon, multi-tool tasks!
- Matches Opus 4.8: Hits 54.4 on DeepSWE, an absolutely insane jump from V4-Proโs previous 8% score!
- Builds Repos from Scratch: Dominates GLM-5.2 (54.2 vs 48.9) on the NL2Repo benchmark.
- Security Beast: 76.7 on CyberGym for heavy-duty security code analysis.
- Fraction of the Cost: DeepSeek locked in a permanent 75% discount on the V4 Pro API. This undercuts GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.7 by 3-19x!
Top-tier frontier performance, but way cheaper.
ahhhhhhh so I'm hiring a tech to help me out in lab and screening applications now and THE KIDS ARE USING PROMPT INJECTION!!!!!! 2.25 pt white text, here's what I've found so far