Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help.
Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written.
Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized.
We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before.
You can read about the process on our Science Blog: https://t.co/ryYnDEAU6J
And see the complete proof on GitHub: https://t.co/wlYMXYnofz
Please welcome to the world a beautiful new geometric object, to do with a problem i’ve always loved. claude really contains multitudes:D Does S^6 admit a complex structure?
Yup
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
Introducing Instaplay: Make or Play Anything
The world is burning trillions on AI to make us more productive - almost none of it has made life more fun. @InstaplayAI is changing that.
100,000+ users are already exploring what games can become when anyone can make one. Creators have already built multiplayer zombie FPS campaigns, interactive classroom lessons, and someone even proposed with a game.
Our mission is to make AI the most powerful force for fun in the world: https://t.co/hZ5kJYAVpT
Introducing Instaplay: Make or Play Anything
The world is burning trillions on AI to make us more productive - almost none of it has made life more fun. @InstaplayAI is changing that.
100,000+ users are already exploring what games can become when anyone can make one. Creators have already built multiplayer zombie FPS campaigns, interactive classroom lessons, and someone even proposed with a game.
Our mission is to make AI the most powerful force for fun in the world: https://t.co/hZ5kJYAVpT
We just raised $5.7M for @PolarBrowser, the AI browser that beats Anthropic and OpenAI on every major web agent benchmark.
- 4,500,000+ actions taken for users, automating sales, recruiting, and ops
- One company cancelled Clay and saves 25+ hrs/wk per person
- Team is from MIT, YC, Prod, Citadel, Jane Street, Perplexity, Modal, and Apple
100 hours of work. From 30 seconds of typing.
Download the world's most powerful AI browser: https://t.co/lUOm01Stm0
hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Training LLMs with NVFP4 is hard because FP4 has so few values that I can fit them all in this post: ±{0, 0.5, 1, 1.5, 2, 3, 4, 6}. But what if I told you that reducing this range even further could actually unlock better training + quantization performance?
Introducing Four Over Six, a new method for improving the accuracy of NVFP4 quantization with Adaptive Block Scaling. 🧵
hillclimb (@hillclimbai) is the human superintelligence community, dedicated to building golden datasets for AGI.
Starting with math, their team of IMO medalists, lean experts, PhDs is designing RL environments for @NousResearch.
Today, we’re announcing Fulcrum Research, a startup scaling human oversight.
We are building debuggers that tell you why your agents fail, and what your rewards are truly testing for—the first step toward the inference-time infrastructure required to safely deploy agents.
Paradigm is the AI-native spreadsheet to eliminate menial work. Thousands of users have saved 10,000+ hours with Paradigm, and you can be next.
Get your first month free today, then plans start at just $20/month.