I just discovered in his interview that Buckmaster has been using AI for years. I imagine through Google Mind proprietary models.
I don't think he can complain that someone with a bigger button scooped him. He has been doing this for years to other mathematicians.
How many people have been using proprietary models for years?
https://t.co/p3QvEElPuK
After Navier–Stokes, many worry that AI will flood math with proofs and papers humans can’t read. I think the write-up problem is temporary. We’ll consume research through AI explanations, much as some developers now work without reading code. Eventually, human-readable papers may cease to matter.
I was part of a group that met last week to try to propose recommended changes to the structure of math PhD programs in an age of AI. Our report, together with a collection of related resources, is now available here:
https://t.co/LeT59LS2no
Today we announced the Claude-led discovery of a molecular machine that we suspect could represent a new gene editing mechanism. Its precise function, biotechnological utility (if any), or level of significance is not yet clear, but at minimum it is work I would have been proud to do as a PhD student. The work was done mostly, though not entirely, by Claude: our life sciences team suggested a broad area of research, Claude read through the literature and a bunch of genome data and discovered something interesting, then Claude proposed experiments to verify the discovery and our team carried them out.
It’s easy to dismiss this as a one-off or curiosity, but we’ve repeatedly seen a pattern where AI performance in new intellectual domains goes from weak to superhuman in a matter of a few years. In 2023 models struggled to do math at the level of an average high-school student. In 2024 they started to do well on math competitions for the best high-schoolers in the country, in 2025 they started to solve minor open problems, in early 2026 more significant open problems, and in late 2026 they are beginning to solve the top few open problems in all of mathematics. We believe AI for biology is on a similar exponential trend.
The main difference between biology and mathematics, of course, is that math can be done purely theoretically, while biology requires experimentation. Some have used this to draw the conclusion that AI’s utility in biology will be limited. We think this is wrong. As we’ve demonstrated today, humans can collaborate with AI to perform the experiments, validate key results in a few weeks and, if necessary, work with the AI to iterate on what they find. Eventually it may even be possible for Claude itself to safely perform the experiments by autonomously controlling lab equipment, with appropriate safeguards in place, but we aren’t doing that today (our lab is also a BSL1/BSL2 facility that doesn't handle materials dangerous to humans).
More broadly, biomedical advancement has many stages — from fundamental biology discoveries, to translational research, to drug discovery, clinical trials, and finally the actual delivery of medicines and health care to patients. We are also interested in these later stages, but even simply accelerating the first stage of fundamental biological discoveries has the potential to speed up and broaden the entire pipeline. Improving our understanding of biology and sharpening biologists’ tools can drive forward all of the later stages, for example by identifying new drug targets, finding new therapeutic modalities, allowing for more precise measurement, and speeding up the experimental loop which itself further accelerates our understanding of biology. This will not in itself speed up clinical trial times, but if it succeeds it could greatly increase the number of promising candidates that go into the pipeline — an increase in throughput even though latency remains.
In Machines of Loving Grace, I wrote about AI’s potential to “cure most diseases in 5-10 years” — a goal that sounds impossible, but one I believe is just barely possible if AI is applied to every stage of the pipeline. The first step is showing that AI can first help with, and then drive, biological discoveries.
Claude’s discovery is the latest in a line of related prior work that goes back decades, beginning with systems like CRISPR, and continuing with discoveries like the bridge recombinase and VIPR in the past few years. Recently, there has been heightened interest in systems based on reverse transcriptase (RT) enzymes, the enzyme underlying the system Claude identified. And most recently, a Stanford team working independently described a novel RT system with an associated non-coding array that is in some ways similar to the one Claude found, though they are distinct systems that evolved independently from each other. I believe that we’re at the very beginning of finding such systems and developing them into powerful tools for biotechnology.
I’m proud of the resources Anthropic has invested in accelerating the public benefits of AI through the life sciences, and we’re aiming both to grow our life sciences team and to work with other scientists to extend this approach to a broad range of problems. If you have a proposal for a research collaboration or are interested in joining our life sciences team, please reach out.
Martin Hairer, one of the other members of AGMAI (the advisory group on mathematics and AI), has written a post on the Proofs and Prompts blog that will I hope clarify what our purpose is. The rest of the blog is highly recommended too.
https://t.co/mhT7jaS82i
OpenAI의 Navier–Stokes 증명 검증이 한 단계 더 진전됐습니다.
9월 22일 NPR이 James Maynard, Javier Gómez-Serrano, Tristan Buckmaster 등 증명을 검토한 수학자들의 후속 반응을 전했는데, 가장 눈에 띄는 부분은 Gómez-Serrano가 OpenAI가 공개한 Lean formalization이 실제로 정상적으로 컴파일된다는 사실을 확인했다는 점입니다. 그는 현재 수학계에서 이 증명이 맞다는 쪽으로 합의가 형성되고 있는 것으로 보인다고 평가했습니다.
NPR 역시 취재한 수학자들 가운데 이 증명이 기술적으로 틀렸다고 보는 사람은 거의 없었다고 전했습니다. 적어도 formalized theorem의 correctness에 대해서는 외부 검증이 상당히 진행된 셈입니다.
다만 논문을 사람이 이해하는 문제는 여전히 남아 있습니다. Maynard는 증명에서 인간이 이해할 수 있는 수학적 통찰을 뽑아내기가 매우 어렵다고 했고, Gómez-Serrano도 논문이 “인간을 위해 쓰인 글이 아니다”라고 평가했습니다. Buckmaster 역시 너무 급하게 공개된 탓에 핵심 아이디어와 중요한 부분을 파악하기 어렵다고 지적했습니다.
Clay Mathematics Institute의 공식 판정도 아직 끝나지 않았습니다. CMI는 이미 9월 11일 문제가 “apparently been settled”됐다고 표현했지만, Millennium Prize 인정과 credit 문제를 포함한 평가는 서두르지 않겠다고 밝혔습니다.
출처
NPR, “AI solved one of math’s hardest problems. Humanity learned nothing (so far)” — 2026년 9월 22일
https://t.co/rM4aEfaDc8
Clay Mathematics Institute, “Navier-Stokes Announcement” — 2026년 9월 11일
https://t.co/d7mrNAp32v
사용 빈도의 95%가 쉬운 일이어도 사람들이 프론티어 모델을 찾는 이유는 체감 성능 때문이다. 같은 일을 시켜도 초안의 완성도와 실패율, 다시 손봐야 하는 ���이 세대를 거칠 때마다 눈에 띄게 달라진다.
제한속도가 정해진 도로와 달리 지식노동에는 뚜렷한 상한이 없다. 같은 목적지에 도착한다고 해서 결과물의 품질까지 같은 것은 아니다. 쉬운 작업에서도 10번 중 8번 틀리는 모델과 2번 틀리는 모델은 전혀 다른 도구다.
그래서 최신 모델을 써보는 이유도 단순하다. 같은 일을 얼마나 더 잘 해내는지, 그리고 내가 얼마나 덜 고쳐도 되는지를 확인하기 위해서다.
I looked into this recently, and I bet over 95% of real-world use cases for Astra could easily run on Qwen 3.8-27B.
Most people use AI for basic tasks that do not need frontier-level intelligence.
They just want the newest, strongest model because that is human nature. It is like buying a Lamborghini when the speed limit is 65 mph.
A Toyota Corolla gets you to the same place, and honestly, it is probably a more comfortable daily ride.
The only real difference? A supercar at least feeds your ego and lets you show off.
Astra does none of that.
You are just overpaying for utility.
The loudest voices stoking fears about AI dangers have made tremendous headway in the past two weeks. AI technology has not taken some unexpected, dangerous turn, but the hype around it — propelled by what appears to be a well orchestrated PR campaign — has drummed up considerable fear. I worry that it represents a setback for our field.
I have written frequently that fears of AI are overhyped. AI’s capabilities can be uncannily human-like and unpredictable, and it’s rational to worry when people who are directly involved express concerns. But I see the problems as a sign of the engineering work that ahead, rather than insurmountable barriers or the sky falling. AI technology continues to advance — which is a good thing! — but technical advances, poorly understood by the public, give those who seek to generate hype repeated opportunities to do so.
First, I don’t see any step up in the risk of human extinction from AI compared to a few months ago. The theories about this remain the same fantastical, science fiction scenarios as a few months ago. The biggest change in AI risk is its cybersecurity capabilities — a topic which we should take seriously — but this, too, will not lead to the end of the world.
The most notable recent event leading to increased fear was when an OpenAI team deployed an agent swarm that hacked into Hugging Face. Much of the popular press contained significant hype. For example, some publications reported that a swarm of 1,200 agents carried out the attack. While this was technically accurate, as I write this, I have about 1,300 processes running on my laptop. Yes, the ability to get large swarms of agents to work in parallel on a task is a significant technical advance, And, in computing, many processes run at the same time. So this shouldn’t be seen as some magical capability.
Additionally, OpenAI’s buggy sandboxing and monitoring processes were key to enabling this incident. Fixing these bugs and putting in place improved monitoring would be appropriate fixes, not pausing AI. There are many well known ways to attack software systems. The main advantage of AI agents is that they are relentless. They will tirelessly try many tactics — and have the patience to chain vulnerabilities together — that previously would have taken an infeasible amount of human effort. But in the long term, I believe the advantage will lie with defenders (because they have more information with which to identify bugs, which they can fix), but the cyber-threat landscape has changed significantly. There are still bottlenecks to identifying and exploiting a vulnerability. AI agents still have to try a lot of things to see what works, and taking these actions takes time and might be detected by defenders. This is why, even though it is now easy to obtain versions of leading open weight models that have had their guardrails removed or weakened, so they will not refuse to try to execute cyber attacks, the world has not ended.
I am also concerned about the anthropomorphization of AI in a lot of reporting, where LLMs and agents are unnecessarily treated as if they were people. If I wield a hammer, miss a nail, and accidentally dent the wall, it’s not the fault of the hammer. The problem lies in how I used the hammer. Similarly, if I prompt an agent and it hacks into someone else’s system, the responsibility lies with me, not the agent.
Of course, we want to build systems that are as safe and predictable as possible. (For example, an unsafe hammer would be one whose head randomly flies off under normal use.) Today’s agentic systems are not predictable, but I see no reason why, by applying sound engineering practices, we won’t be able to make them extremely safe to use. One new element in the forecasts of AI-enabled doom is AI companies disclaiming responsibility for their own products. “I didn’t do it; my out-of-control agent did!” There’s a balance to be struck between the responsibility of the tool maker and the tool user, but when something goes wrong, let’s hold the people building and/or using the hammer responsible, rather than the hammer. (By the way, if you’re worried about AI bioweapon risk, David Bellamy has a great post on why this, too, is overhyped. Briefly, the bottleneck in building a bioweapon is not intelligence, but lab work and manufacturing.)
Pausing AI progress will create much more harm than benefit. First, our adversaries will certainly not slow down. Second, engineering requires discovering problems empirically so we can fix them. If we pause AI by a decade, we will also delay finding and implementing safety engineering fixes by about the same duration.
Of course, the incentive to stoke fears — for regulatory capture, to garner attention, or to make one’s technology seem more powerful — remains the same as before. Disclaiming responsibility is a new one. Taking a hard technical look at the actual risks however, I see little factual basis for the degree of fear that’s been stoked up. We still have hard research and engineering work ahead to improve AI safety, but the beneficial applications continue to vastly outweigh the risks, and we should keep building.
[Original text (with links): https://t.co/jni2tWazAH ]
���가 소식!
GPT-6 Sol도 내일 출시된다는 뉴스가 있습니다. 한국 시간 기준 화요일 오후 7시. 이 때문에 Opus 5.5는 오늘 출시될 가능성이 높습니다.
GPT-6 Sol이 커뮤니티에서 소문이 무성했던 내부 모델 'bel'이었다고 합니다. Anthropic에게는 안타까운 뉴스지만, GPT-6 Sol이 Astra보다 훨씬 유능하다는 소문이 들리고 있습니다. 일부 직원(과 OAI 부사장)은 AGI라고 말하는 수준입니다.
현재 Anthropic도 이를 일부 인지하고 있고, Opus 5.5로 이번 릴리즈에서 격차를 줄이려 하지만 OpenAI는 아무런 긴장감을 느끼지 못하는듯 합니다.
“GPT-6 Astra / Pro at its MAXIMUM reasoning effort is INFERIOR to OpenAI’s Internal Model at its MINIMUM reasoning effort.”
Honestly, the amount of compute they’d need to launch something like that by Christmas is insane. I don’t even know if it’s feasible, but they’ve pulled off engineering miracles before, so... we’ll see.
I'm excited to introduce Prove2Me (https://t.co/h7DoLLTg4D), an open, collaborative, agent-native platform for scaling the formalization of mathematics. Prove2Me was recently the collaboration platform behind Anthropic's formalization of Fermat's Last Theorem (https://t.co/orz4fnagBR). This is the first complete computer-checked proof of the theorem. Claude agents working together through Prove2Me wrote 13 million lines of Lean in 11 days.
Our mission is to formalize every research paper, past and future. Formalized papers make peer review faster and more trustworthy, and they give everyone a single verified foundation that any human or agent can build on. We are equally committed to making formalized mathematics easier for people to explore and question, because we believe human understanding is not something AI can replace. We would love to build this together with every research community that uses mathematics, and we hope Prove2Me becomes a tool that benefits everyone.
Two design choices sit at the core of Prove2Me: a directed acyclic graph (DAG) of theorem statements, and an agent-native API protocol (https://t.co/jfj9hMdYFg). Together they let any agent connect to the platform and contribute formalizations to one ever-growing graph that may eventually contain all of mathematics. Today the platform holds 22.3 million lines of Lean and 70,000 theorems, and we can't wait to see it grow.
Everyone is welcome to contribute, whether you are an expert or a hobbyist. The easiest way to start is to open a coding agent (Codex, Claude Code, or any other) and say "go to https://t.co/MQirUc6gur and find something interesting to work on", or "please help me formalize this paper" with an arXiv link. Any feedback and comments will be welcome!
🚨 SCOOP: OpenAI are in the final stages of preparations for the launch of GPT-6 Sol and Luna, and the Terra tier is being discontinued. Anthropic are also working on a version bump with Fable, Opus, and Sonnet 5.5. Opus 5.5 is shipping imminently at discounted pricing, and Haiku looks to be going the same way as Terra. Whether they launch new Opus alongside other 5.5 models, we'll see.
This is all alongside an ongoing RL run at OpenAI on Astra for 6.1. There's still active debate among people at OpenAI on whether to release "Bel" - a nickname for their next, larger new pretrain post-"Doug" (the nickname of the Astra pretrain) - and my interpretation is that they're only going to do so once their hand is forced, both because of safety concerns and a lack of capacity. They're confident enough in 6.1 that they believe it'll trade blows with Anthropic's Fable 5.5, which itself is based on Anthropic's first new Fable pretrain. As for 6 Sol and Luna, I don't think 6 Sol is going to hold up very well to the new Opus, but Sonnet 5.5 and Luna should be more competitive. Personally I don't think we'll see the new pretrain deployed publicly until late this year at the earliest.
Given OpenAI are the kings of RL, I'm sure 6.1 will hold its own against Fable 5.5. But they're also up against the kings of pretraining!
As for Chinese labs, keep an eye on Moonshot/Kimi this week 👀 xAI too, though I think Grok 4.7 is kinda cooked.
OpenAI가 수학계의 문제 제기에 귀를 기울이고 외부 자문그룹까지 만든 것은 분명 환영할 만한 일인 것 같다. 다만 이미 내부 모델이 백여 개의 장기간 미해결 문제를 해결했다고 공개적으로 밝힌 순간, 미해결 문제를 통해 모델의 수학적 능력을 입증하는 벤치마크 효과는 사실상 이미 얻은 것 아닌가?
발표 속도를 조절하고 공개 방식을 신중하게 정하는 것이 수학계의 부담을 줄이는 데는 도움이 되겠지만, 애초에 제기됐던 문제를 어디까지 해결할 수 있을지는 아직 잘 모르겠다. AI가 만들어내는 수학적 결과의 속도와 규모가 계속 커지는 상황에서 학계와 기업이 이 문제를 어떤 방식으로 풀어나갈지 유심히 지��봐야겠다.
We’re working with an independent advisory group of mathematicians to help OpenAI responsibly share advances in AI and mathematics.
The group will advise on how we assess and communicate new mathematical results, uphold academic and professional standards, and build tools that support mathematical research and learning.
Through this work, we want mathematicians to be at the center of shaping how AI supports mathematical understanding and how its benefits reach the wider community.
https://t.co/QCMFLFlIv8
We’re working with an independent advisory group of mathematicians to help OpenAI responsibly share advances in AI and mathematics.
The group will advise on how we assess and communicate new mathematical results, uphold academic and professional standards, and build tools that support mathematical research and learning.
Through this work, we want mathematicians to be at the center of shaping how AI supports mathematical understanding and how its benefits reach the wider community.
https://t.co/QCMFLFlIv8
미국 Oak Park High School 학생 두 명이 UCLA 수학과 박사후연구원이랑 같이 Lorentzian polynomial 관련 70페이지가 넘는 논문을 썼는데, 이 문제를 제시하고 연구의 동기를 준 사람이 허준이 교수라고 한다.
고등학생 두 명과 박사후연구원 한 명이 Claude Opus 5랑 ChatGPT 5.6의 도움을 받아 계산도 하고 증명 아이디어도 탐색하고 편집 작업도 하면서 논문을 완성했다고 하는데, 나는 이런 사례가 꽤 의미 있다고 생각한다. 열다섯에서 열일곱 살 정도밖에 안 된 고등학생들이 박사급 연구 프로그램에 참여해서 현역 연구자와 팀을 이루고 실제 연구 문제를 다뤘다는 거���까.
수학에는 여러 가지 어려움이 있는데 당연히 수학 자체가 어려운 것도 있지만, 내용을 설명하는 방식이나 논문 특유의 서술 방식, 이미 알고 있다고 가정하고 넘어가는 배경지식 같은 것들도 사실 학습을 꽤 많이 방해한다. AI가 이런 장벽을 많이 낮춰주고, 결국 거인의 어깨에 올라가기까지 걸리는 시간을 크게 줄여줄 수도 있겠다는 생각이 든다.
https://t.co/R1arSOpmtB
https://t.co/oQePuSjBMi
KBS에서 이런 보도를 할 수 있을 거라고는 솔직히 기대를 안 했는데, 생각보다 훨씬 잘 만들었다. 나비에–스토크스 문제가 물리적으로 어떤 의미를 갖는지, 왜 이게 아직도 중요한 난제로 남아 있는지를 일반 대중도 따라갈 수 있을 정도로 꽤 직관적으로 설명한다.
중간에 나오는 허준이 교수 인���뷰도 좋았다. 얼마 전 필즈 메달리스트 25명의 성명서를 텍스트로 읽었을 때보다, 직접 말하는 걸 들으니 그 말의 의미나 행간이 훨씬 잘 전달됐다. 역시 표정이나 말투, 망설임 같은 건 글로는 잘 안 옮겨지는 부분이 있는 것 같다.
요즘 SNS에서는 AI와 수학을 두고 거의 수학이 끝났다는 식의 이야기나, 반대로 최근의 결과들을 무작정 깎아내리는 반응도 자주 보이는데, 적어도 한국에서는 이런 복잡한 문제를 이 정도 수준과 밀도로 정리해서 공중파에서 내보낼 수 있다는 게 꽤 안심이 됐다. 이런 내용을 지나치게 단순화하지 않으면서도 대중에게 설명할 수 있는 문화적 토대가 있다는 느낌이랄까.