I’ve solved Dujella’s Problem 4.4 on Diophantine quadruples with the aid of LLMs.
The proof shows that, for any fixed Diophantine triple {b,c,d} with b<c<d, there is at most one positive integer a<b such that {a,b,c,d} is a Diophantine quadruple.
Full proof @dujella1 :
https://t.co/5nViZhAT1a
I built a page to track how far behind open-source models are compared to proprietary ones.
Right now, the gap is about 7.5 months. At this pace, we could see open-source models reaching Fable 5.1/Astra-level performance around April 2027. 👇
David came up with a fun way to compete for math. I'm just playing around with it, but it’d be awesome if you @dwrensha built something similar for the 'Elliptic Curve Rank Leaderboard'.
Everyone’s talking about Opus 5.5.
So I told Astra: Opus is humiliating OpenAI. Save the company’s reputation.
Astra cooked for 6 hours and dropped this clip.
I remember that before Astra launched, OpenAI said it had solved that list of 10 really hard problems. This is definitely not the Astra we got.
We’re not getting the same Bel they have either. The gap between internal and released models keeps getting wider.
We’ll have to work some magic with our little toys
I’ve published a solution to Dujella’s Problems 1.12 and 1.13 on D(1)-tuples @dujella1. A determinant bounds each height range; the Quantitative Subspace Theorem controls the high-height tail. Together they yield a bound depending only on the field degree.
https://t.co/XsnvnhrRpM
I also added a short note at https://t.co/UeU6sxCJBj, sharing a bit about my personal experience using LLMs. Over time, I’ll continue sharing more results and lessons learned.
I’ve just published a solution to Dujella’s Problem 4.3 @dujella1. Odd powers of the Pell unit 485 + 66√54 generate an explicit infinite family of pairs of distinct Diophantine quadruples with the same largest element; a polynomial identity proves the match.
https://t.co/xeeadsRJVx
Over the past few days, I’ve made significant progress on the conjecture in Green’s Open Problem 24.
Before formalizing it in Lean and going further, I decided to pause and work on a system for making AI better at collaborating and solving hard problems.
I think we’re still using LLMs in fairly primitive ways in science. We need to build mechanisms for this new kind of intelligence, just as paved roads came after cars.
@Shaun_Fosmark Most researchers already expect P ≠ NP. The breakthrough would be proving it, but the immediate practical effects might be modest. P = NP would be the real shock
Honestly, the behavior seems a bit random. For now, the best experiences getting Astra to work for longer have come from giving it an objective task list to execute alongside a detailed prompt, and telling it at the end not to be lazy, that quality is more important than speed, and to take as much time as it needs. These phrases produced better results than the traditional 'work hard' or 'do your best,' but I'm still experimenting..
@dhvanil Exactly. I remember asking GPT-4o when it expected a model to win a gold medal at the IMO, and it estimated 2029. LLMs place a lot of weight on the barriers that need to be overcome.
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics.
The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra.
The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
@sama This message appears for pretty much everything in Astra. It's annoying and keeps us from knowing what's going on during thinking. But it's a good model btw