@PatrickKidger@EdwardRaffML If any startup, manager, or company expects you to work these hours its a massive red flag. There is zero evidence that this leads to any long term productivity. Burn out is real.
This conjecture was, by far, the most interesting open problem in my area of statistics. Any statistician would have been excited to construct this counterexample, and I’d have been thrilled for them.
GPT-5.6 solved it, but I wish a human had.
AlphaFold3 is no longer the best model.
Over the past year, we've seen many (mostly companies) claim that their newly released models matched or outperformed AlphaFold3 in some important structure prediction task. A substantive amount of these also underperformed their self-professed benchmarks under independent scrutiny.
This is why I was initially skeptical on the results from the ESMFold2, Protenix-v2, and OpenDDE teams' respective releases. Within a few weeks of each others, these groups each released benchmarks showing meaningful outperformance over AF3.
Finally, it is a good moment to talk about the hardcore math details of the work that GPT 5.6 helped me with.
The whole journey started in 2023 when my colleague @PiotrPePo emailed me with a question about algebraic K3 surface. The idea was that he was looking for configurations of lines on such surfaces which have remarkable property, namely that they have unusually high logarithmic Chern slope. In laymen terms, such objects are extremely rare and highlight a very special type of geometry on the surface.
It was known for many years that due to Bogomolov-Miayoka-Yau inequality such a slope cannot be larger than 3 and for very symmetric quartic and some other classes of algebraic surfaces it was known to be at most 8/3.
In our preliminary work together, we proved that the famous Fermat quartic x^4+y^4+z^4+w^4=0 (projective coordinates) contains a subconfiguration of lines with slope 8/3. This was a discovery which I quickly made after our initial exchange of emails with Piotr. Since 2023 we started thinking about a possible way to prove this conjecture that perhaphs, since Fermat's quartic is so symmetric, it might be the actual maximum.
We have published a very nice survey paper with new results in EMS Surveys in 2026
https://t.co/zLSl7sBbMQ
Initially, in 2024 and 2025 I've been asking various AI models to find some kind of hint for us, whether there is a way to make and organize our computations in a better way. An initial hint came from GPT Pro to use systematically all the inequalities that we've been using and turn this whole venture into some kind of linear programming computation.
This initial hint was partially interesting but we could not get past 8/3 with the techniques that we had.
But in the paper we had a lot of nice examples of surfaces which contained billions of different subconfigurations of lines. So a brute force search was basically impossible.
And here comes GPT 5.6 Sol which I used to reiterate on this problem. The model, prompted with many pages of my technical documentation, in Codex, managed to find a critical paper about fractiona linear programming, an idea from a paper
W. Dinkelbach, On nonlinear fractional programming, Management Sci. 13 (1967)
The model immediatelly implemented the whole setup and tested it against the many families of K3 surfaces that we collected.
It was a big shock to me when after basically about 30 minutes of work, around 4 am in the morning I finally saw a counterexample to the slope conjecture. There exists a configuration of 24 lines on Schur quartic which provides the slope 14/5. Initially, I could not believe that this was true but I quickly prototyped a suitable prompt and code which verified everything perfectly. From this insight we understood that we were completely wrong about the conjecture because we have assumed that the maximum should come from the most symmetric K3 surface, but instead the right idea was to look for the most symmetric configuration!
After all, I think this whole experience made me realize how important it has become in my scientific life to use high quality models. We could have spent another hundreds of hours trying to prove a wrong conjecture. Now we finally understand what is possible for K3 quartics. The answer is a major advancement in our understanding of algberaic geometry of surfaces and shows that a fruitful collaboration between humans and AI is possible and can empower humans to get better insights into mathematics.
The new paper will appear tonight on arXiv.
Are therapeutics startups the crabs of the biotech world?
In evolutionary biology, there's a phenomenon called carcinization, where different species, again and again, independently evolve into something that looks like a crab.
It's happened at least 5 times.
🧵 (11)
Do single-cell foundation models obey scaling laws?
A somewhat thought-provoking new Nature Methods study by the Crawford lab suggests that, for current single-cell foundation models, the answer may be “not really.” Across a broad range of architectures and downstream tasks, increasing pretraining data from hundreds of thousands to tens of millions of cells yielded surprisingly limited gains, with performance often saturating much earlier than expected.
This is interesting and provides exactly the kind of rigorous benchmarking our field needs. As Felix Fischer and I commented in the accompanying Research Briefing, such studies help move the discussion beyond model size and computational budgets toward actual scientific utility.
At the same time, I am not convinced the key conclusion is that scaling does not work in biology. Rather, it may be that current objectives are not extracting enough information from additional data.
Interestingly, in our recent scConcept work, we observe a markedly different scaling behavior, with continued gains as training data grows toward hundreds of millions of cells. The key difference may be the training objective itself: instead of reconstruction-based masked modeling, scConcept uses a contrastive objective that directly optimizes biologically meaningful cell representations.
https://t.co/CR8DSUmHi3
This raises an interesting question for the field: Have we reached the limits of data scaling, or only the limits of current objectives?
-> My guess is that the next generation of biological foundation models will depend less on simply collecting more cells and more on finding the right representation learning principles for biology.
Nature Methods paper:
https://t.co/ArtUo1fcPE
Research Briefing:
https://t.co/5er4vGJiAJ
#SingleCell #FoundationModels #AIforBiology
Over lunch today I was asking Claude to help me feel less disillusioned by the direction of modern biology, and it gave me a well-reasoned reply that did precisely the opposite.
Claude nails it here:
Today I'm posting a new blog about an astonishing scientific own goal. Hundreds of papers have reported using a completely wrong antibody to investigate the tumor suppressor p16. This mistake has happened because scientists have muddled the names of two proteins 🧵
Following up on the suggestion from Will Sawin, here is an illustration of the new configurations that disprove Erdos' unit distance conjecture (made with the help of ChatGPT 5.5 Thinking).
I propose a life ban from arXiv when there is an argmin/min/argmax/max/expectation operator without saying over **what**, or when the "over what" variable doesn't appear in the operand expression.
Some concrete examples here:
👉 modern cardiovascular drug development has been largely shaped by targets, where variation in their genes has been linked to disease risk (PCSK9, APOC3, ANGPTL3, FXI, LPA, IL6 etc.)
👉 specifically, human genetic evidence has been translated to approved drugs for hypercholesterolemia (PCSK9, approved), boosted the clinical development of new categories of drugs (e.g. LPA, IL6, in phase 3), or informed new translational pathways and indications for others (e.g. FXI, awaiting approval for non-cardioembolic ischemic stroke)
👉 GWAS hits in Alzheimer's disease (e.g. TREM2, CLU, BIN1) have pointed to neuroinflammation and microglial biology, influencing drug development, even if clinical translation is still ongoing (TREM2)
👉 similar story for ALS, where beyond pointing to Mendelian cases, for which targeted therapies are being developed, drugs emerging from GWAS hits are about to start being tested in clinical trials (UNC13A)
👉 many similar examples from autoimmune diseases, e.g. IL23/IL23R approved for psoriasis and inflammatory bowel disease, TYK2 inhibition approved for psoriasis now expanding to other indications, PTPN22/LYP emerging as a promising target in clinical development
👉 first approved drugs for dry age-related macular degeneration (C3 and C5 inhibitors) targeting the complement, were largely influenced by discovery of hits in several complement genes in GWAS
👉 the whole booming field of RNA therapeutics (especially liver-targeting) depends on sequencing the targets and so is directly linked to the human genome project