A tectonic shift in mathematical discovery is unfolding before us. AI is advancing on problems that resisted decades of work, and it doesn’t have to cost millions. Our agent made the first improvement in 38 years on a central problem in extremal graph theory for under $2,000.
When I started my master’s in 2014, we were chasing better scores on CIFAR-10 and ImageNet. It’s hard to wrap my head around how far we’ve come.
1/ Double Agent, our agentic system, has achieved a breakthrough result on one of the most central problems in extremal graph theory. It constructed explicit regular graphs with girth (8/5) log n, the first improvement in 38 years, and a leap towards the theoretical maximum.
1/4
On a hard problem, an AI's first answer is rarely its best. The gains come from generating many tries and keeping the best. But standard RL collapses a model onto one answer, so more samples don't help. ArgMaxRL fixes that for continuous rewards, not just pass-or-fail.
Quantum circuits. GPU kernels. pandas. Chip design.
Different domains. One system: WarpSpeed.
We updated our research page. The blogs, the papers, and the benchmarks now sit in one place.
https://t.co/Sc70hVTIVW
Our agent has this slightly ridiculous habit of serendipitously solving whatever domain we point it at: GPU kernels, CPU code optimization, then cryptography-breaking quantum circuits, and NOW chip design. It keeps surprising me how well it generalizes across fields.
1/5
Designing a chip can take years, hundreds of specialists, and tens of millions of dollars.
WarpSpeed, our AI expert, just reached state-of-the-art performance on @NVIDIA's CVDP benchmark for chip design, surpassing Claude Code with Fable.
We built WarpSpeed to solve really hard problems, starting with code optimization. But I guess when you have a really good hammer, a surprising number of things turn out to be nails. 🔨
1/ No one is breaking Bitcoin any time soon.
But the projected cost of the attack just dropped. WarpSpeed, our AI expert, found a circuit 2.5x more efficient than the one Google's quantum team published in March.
1/ No one is breaking Bitcoin any time soon.
But the projected cost of the attack just dropped. WarpSpeed, our AI expert, found a circuit 2.5x more efficient than the one Google's quantum team published in March.
Writing fast, correct kernels is very hard! For both speed and logic, verification is surprisingly tricky.
Very cool work, and writeup, from @_doubleAI_
https://t.co/vuyVEtpUmN
Interesting new result from @_doubleAI_ : their analysis of Nvidia’s SOLBench shows that agentic frameworks were overfitting to the benchmark environment on some of the problems. DoubleAI’s approach not only reduce overfitting but also substantially improves performance, running 2.24x faster on average. This work offers an early glimpse of what a future “theory of generalization” for agentic systems might look like.
WarpSpeed, our autonomous optimization agent at @doubleAI, just took first place on @NVIDIA's new SOL-ExecBench: 235 of the hardest CUDA kernels in production.
But the more interesting story is what we found along the way.
Verifiers designed for human errors don't defend against AI reward hacks. We found four ways the same benchmark's verifiers can be silently fooled. The first one broke transformer training.
🧵
Not all benchmarks are made equal.
@NVIDIA’s SOL-ExecBench uses real AI training and inference kernels.
In AI, compute is the bottleneck. Even 10% faster matters.
Our kernel optimization agent worked for a single day and hit #1: 2.24x avg speedup.
So proud of @_doubleAI_.
We ran WarpSpeed, our autonomous optimization agent, on @NVIDIA's new SOL-ExecBench for a single day.
It took first place by a wide margin, beating the optimized kernels on 90% of problems, with an average speedup of 2.24x.
ExecBench gathers 235 of the hardest CUDA kernels in production today, lifted from real workloads in DeepSeek, Qwen, Gemma and Kimi.
Blackwell kernels are notoriously hard to write. But we find that verification is just as hard.
We have a story to tell.
https://t.co/aQX9XXCz4z
We ran WarpSpeed, our autonomous optimization agent, on @NVIDIA's new SOL-ExecBench for a single day.
It took first place by a wide margin, beating the optimized kernels on 90% of problems, with an average speedup of 2.24x.
ExecBench gathers 235 of the hardest CUDA kernels in production today, lifted from real workloads in DeepSeek, Qwen, Gemma and Kimi.
Blackwell kernels are notoriously hard to write. But we find that verification is just as hard.
We have a story to tell.
https://t.co/aQX9XXCz4z
החברה המסתורית של אמנון שעשוע ושי שלו-שוורץ, doubleAI, גייסה מאות מליוני דולרים (והוקמה קצת על הראש של AI21 שמחכה כנראה על המדף לרוכשים). עכשיו הדלת של המפעל נפתחה וסוף סוף יצא משהו, אבל היה שווה לחכות.
למה זה מרשים? החוקרים בחרו בכוונה להכנס באחת הבעיות הקשות ביותר: לכתוב גרסה מותאמת ומהירה יותר של cuGraph של NVIDIA. מעטים המפתחים שיעזו לפתוח את מכסה המנוע של CUDA, ורק הקשוחים שביניהם יתעסקו בגראפים. את הספרייה מתחזקת קבוצה של מומחים מאנבידיה שאינם פראיירים, אולי החבורה שמבינה יותר מכל אחד בעולם את הנושא המורכב (האם חלק מהצוות של AAI מגיע משם ב��קור? לא יודע).
האלגוריתמים שכעת בספריה בנויים באופן כללי יחסית, כך שיוכלו להתאים לקשת רחבה של GPUs וקונפיגורציות. אבל המערכת של AAI חרוצה הרבה יותר. במקום להסתפק במשהו כללי, היא תפרה אלגוריתם ספציפי מיוחד עבור כל צירוף של קונפיגורציות (סה"כ 576 במספר!).
חבורת מהנדסים אנושית לא היתה בוחרת בדרך הזו אף פעם. זו עבודה משוגעת וגם סיוט לתחזק. מקסימום היו עושים אחד או שניים כאלה עבור היוזקייסים הפופולריים ביותר. אז כך, במקום לנסות ליצור ספרייה *כללית* טובה יותר, משתמשים בחוזקה של ה-AI שלא משתעממת לעולם כדי לשפר ספציפית כל מקרה בפני עצמו (שיפורים יפים ומשמעותיים, לא סתם אינסטלציה).
כל העבודה מורכבת ומרשימה מאוד. מרגישה כמו חליפה שנתפרה ע"י חייט עילית. הם בנו מערכת שלמה כדי לבדוק אם התוצאות של המנוע שלהם מדוייקות (כי המערכות הקיימות לא היו מספיק טובות מכמה סיבות.. תקצר היריעה). אימנו את המנוע שלהם על קשת של בעיות מגוונות שהוא המציא לעצמו, וכל קרנל שהוא כתב שימש כמדרגה לשלב הבא.
התוצאה היא שיפור של 3.6x על פני כל האלגוריתמים (חמישית מהירים פי 10!) וכולם כולם 100% מדוייקים ונכונים, כולל בעיות שהם גילו בספרייה של אנבידיה ותיקנו גם.
אם יש פה חיסרון הוא שבערך חצי מהמערכת (נראה לי) בנוי ספציפית לבעייה המיוחדת הזו. כדי להפנות את המערכת לתקוף בעייה חדשה צריך לבנות מחדש, עם הבנה עמוקה של הבעייה החדשה, את המערכת שבודקת ומדרגת את העבודה שהמו��ל יוצר.
אז זה לא משהו שתגידו לו "הי קלוד הנה לינק, לך תשפר את כל מה שיש שם" וזה פשוט יעבוד. כאמור נעשתה פה הרבה עבודה מורכבת וספציפית.
אבל עדיין, יש לנו פה AI ששם לכם בגיטהאב קוד ש��ץ פי 3.6 מהר יותר מזה של מהנדסים מהטובים בעולם שמתעסקים בבעייה שהם הכי טובים בה.
יהיה מעניין לראות אותם תוקפים בעייה יותר מסחרית ופחות אקזוטית. אם ישיגו כזה שיפור לטרנספורמרים, או לדחיסת וידאו, או קריפטוגרפיה או משהו בביולוגיה לא יודע, זה יכול להיות מתורגם להמון כסף (או שאולי יש הרבה כסף באופטימיזציית גרפים?).
דבר יפה מאוד 🫡
I'm really excited by using AI to write kernels and accelerate kernel development. At @togethercompute we use AI models extensively through all parts of the performance engineering pipeline.
Super excited to see this new work coming from DoubleAI!
Some publish benchmarks. We shipped doubleGraph: a drop-in replacement for a widely used GPU-accelerated graph library, rebuilt by our WarpSpeed agent — 3.6× avg speedup, 100% of kernels faster — surpassing a decade of expert GPU engineering. ⚡️
1/ Software was eating the world - and now AI is eating software.
AI already beats humans at math/coding (IMO, CodeForces). Right?
So let's test the strongest coding agents on a real domain: optimizing cuGraph (GPU graph analytics kernels).
Spoiler:
* The strongest coding agents crash.
* And @_doubleAI_ built WarpSpeed - an AI that beat a decade of expert-engineered GPU kernels.
🧵
DoubleAI’s AI system just beat a decade of expert GPU engineering
WarpSpeed just beat a decade of expert-engineered GPU kernels — every single one of them.
cuGraph is one of the most widely used GPU-accelerated libraries in the world. It spans dozens of graph algorithms, each written and continuously refined by some of the world’s top performance engineers.
@_doubleAI_'s WarpSpeed autonomously rewrote and re-optimized these kernels across three GPU architectures (A100, L4, A10G). Today, we released the hyper-optimized version on GitHub — install it with no change to your code.
The numbers: - 3.6x average speedup over human experts - 100% of kernels benefit from speedup - 55% see more than 2x improvement.
But hasn’t AI already achieved expert-level status — winning gold medals at IMO, outperforming top programmers on CodeForces? Not quite. Those wins share three hidden crutches: abundant training data, trivial validation, and short reasoning chains. Where all three hold, today’s AI shines. Remove any one of them and it falls apart (as Shai Shalev Shwartz wrote in his post).
GPU performance engineering breaks all three. Data is scarce. Correctness is hard to validate. And performance comes from a long chain of interacting choices — memory layout, warp behavior, caching, scheduling, graph structure. Even state-of-the-art agents like Claude Code, Codex, and Gemini CLI fail dramatically here, often producing incorrect implementations even when handed cuGraph’s own test suite.
Scaling alone can’t break this barrier. It took new algorithmic ideas — our Diligent framework for learning from extremely small datasets, our PAC-reasoning methodology for verification when ground truth isn’t available, and novel agentic search structures for navigating deep decision chains.
This is the beginning of Artificial Expert Intelligence (AEI) — not AGI, but something the world needs more: systems that reliably surpass human experts in the domains where expertise is rarest, slowest, and most valuable.
If AI can surpass the world’s best GPU engineers, which domain falls next?
For the full blog: https://t.co/sCF033hb28
CuGraph:
https://t.co/jqxrcuhfs4
Winning Gold at IMO 2025:
https://t.co/fAdIT2mTkI
Codeforces benchmarks:
https://t.co/UhRAUieWFi
@shai_s_shwartz post:
https://t.co/1WAGIXfiqh
From Reasoning to Super-Intelligence: A Search-Theoretic Perspective
https://t.co/iX625p57NT
Artificial Expert Intelligence through PAC-reasoning
https://t.co/Hq3wWsmidw
DoubleAI’s AI system just beat a decade of expert GPU engineering
WarpSpeed just beat a decade of expert-engineered GPU kernels — every single one of them.
cuGraph is one of the most widely used GPU-accelerated libraries in the world. It spans dozens of graph algorithms, each written and continuously refined by some of the world’s top performance engineers.
@_doubleAI_'s WarpSpeed autonomously rewrote and re-optimized these kernels across three GPU architectures (A100, L4, A10G). Today, we released the hyper-optimized version on GitHub — install it with no change to your code.
The numbers: - 3.6x average speedup over human experts - 100% of kernels benefit from speedup - 55% see more than 2x improvement.
But hasn’t AI already achieved expert-level status — winning gold medals at IMO, outperforming top programmers on CodeForces? Not quite. Those wins share three hidden crutches: abundant training data, trivial validation, and short reasoning chains. Where all three hold, today’s AI shines. Remove any one of them and it falls apart (as Shai Shalev Shwartz wrote in his post).
GPU performance engineering breaks all three. Data is scarce. Correctness is hard to validate. And performance comes from a long chain of interacting choices — memory layout, warp behavior, caching, scheduling, graph structure. Even state-of-the-art agents like Claude Code, Codex, and Gemini CLI fail dramatically here, often producing incorrect implementations even when handed cuGraph’s own test suite.
Scaling alone can’t break this barrier. It took new algorithmic ideas — our Diligent framework for learning from extremely small datasets, our PAC-reasoning methodology for verification when ground truth isn’t available, and novel agentic search structures for navigating deep decision chains.
This is the beginning of Artificial Expert Intelligence (AEI) — not AGI, but something the world needs more: systems that reliably surpass human experts in the domains where expertise is rarest, slowest, and most valuable.
If AI can surpass the world’s best GPU engineers, which domain falls next?
For the full blog: https://t.co/sCF033hb28
CuGraph:
https://t.co/jqxrcuhfs4
Winning Gold at IMO 2025:
https://t.co/fAdIT2mTkI
Codeforces benchmarks:
https://t.co/UhRAUieWFi
@shai_s_shwartz post:
https://t.co/1WAGIXfiqh
From Reasoning to Super-Intelligence: A Search-Theoretic Perspective
https://t.co/iX625p57NT
Artificial Expert Intelligence through PAC-reasoning
https://t.co/Hq3wWsmidw
Deep reasoning is beyond the capabilities of today’s AI models. GPT5 shows some progress but overall the performance is a far cry to what is required to solve problems at expert level. Statements about models reaching PhD level should be taken with a measure of skepticism.
3/
The problems in FormulaOne are succinct, only consisting of a sentence or two that can be understood by any undergrad, but their solution requires ingenuity and deep reasoning. To illustrate, here are two variations on a problem.
Easy version: A dominating set in a graph G = (V,E) is a set of vertices such that all vertices in the graph are either in the set, or have a neighbour in the set.
Given a graph G, count the number of dominating sets of G.
Harder version: Same as the easy version, but now count only dominating sets of G that contain no other dominating set as a proper subset.