Similar experiences asking Claude/ChatGPT to explain proofs in some Floer homotopy theory papers. Current consumer LLMs are appalling at real research level maths..
It's illuminating that the frontier models are exactly as bad at thinking about geometric topology as geometric topologists are at putting their thought to words.
Here is GPT 5.5 Pro's proof of the extremely well-understood h-cobordism theorem from the 60s. Glaring issues here.
One pretty entertaining detail is that Claude pattern matched on symplectic geometry papers citing the non-compactness of moduli spaces as the main obstacle in Floer theory, which is obviously true in general. But it’s compteley irrelevant in the specific case I pointed Claude to.
Congrats and thanks to Ciprian and Nick on their paper https://t.co/RdHLkjnAId, based on a completely autonomous Aletheia solution, which resolves a Kirby problem from K3! This Kirby problem list was assembled by leading experts in low-dimensional topology, looking specifically for hard and important problems to guide the future of the field. (In this way it differs from the Erdős problem list, whose motivation is historical, although of course there are many important and highly influential Erdős problems as well.)
The Quanta article https://t.co/wLQnUX5n2S gives a lot more background and context about this K3 list. It quotes Maggie Miller and Inanc Baykur, two of the experts involved in creating the list, as saying the following about the problem selection:
>>> “It should be a problem that is sufficiently interesting that if a solution came out, it would have the potential to change the field,” Miller said. Baykur added: “Maybe a small percentage can be solved in the next two to three years.”
That being said, not all conjectures turn out to be as influential as hoped, and Manolescu and Rozenblyum are careful not to overstate the importance of this solution. The problem turned out to be accessible using existing results from the literature, but in areas a bit distant from topology (e.g., decidability theory), which may explain why it had eluded topologists.
Mathematicians telling you math is dying are like boomers telling you civilization is over. They are talking about their own personal obsolescence, not the future. We are barely out of the Neolithic and science is just starting.
Survey time: If you see a cryptographer saying "We told you for many years to trust cryptosystem X, but, oops, it's actually giving away all your data; please panic right now", how much do you believe the same cryptographer saying "We're telling you to trust cryptosystem Y"?
Had to tell ChatGPT three times with growing evidence when it dismissed my claims as fake news until it admits failures. Feels like Petey’s last line of The Birthday Party by Harold Pinter.
Finally figured out how to get ChatGPT to stop BS and pattern matching with fake risk aversion and status bias - just tell it to give a list of lies/fabrications from the immediately preceding exchanges. Worked unbelievably well. Also worked damn well on Claude.
All the sum-check optimizations we could come up with (some brand new from new coauthor Zach DeStefano), collected in one paper.
These make a huge difference in Jolt (~11x faster, ~17x less memory for key subroutines), and can benefit almost all sum-check-based SNARKs.
Testing Claude Opus 4.6 extended thinking about maths questions is the most painful thing I did this week - it was fine on stuff like convex optimization, but for stuff like p-adic modular forms, the only thing it extends is confusion, deception, retraction, eventually apology (only after sustained reproach) - in that order.