It was a real honor to deliver the Opening Keynote at the Royal Society’s AI & Law Conference last week in London ! Thanks to all of the organizers and fellow speakers for this important dialogue regarding the future of our field. Link to Event Here:
https://t.co/JcfqrS866P
NFL legend Mike Ditka, who revolutionized the tight end position before becoming the head coach of the world champion 1985 Bears and a towering figure in the city of Chicago, has passed away, his family announced. He was 86.
What should an open scientific paper look like in the AI age?
We made a demonstration project out of our paper on the childhood socioeconomic circumstances of the Nobel Laureates.
Please check it out and share. Details in 🧵 1/
*Mathematicians did not ask for this work to be done"
WTF?!! Since when was math something that was only done at the request of professional mathematicians?!!
Mathematicians don't own mathematics! This is crazy!!
This got me curious so I went and found the press release which I highly recommend checking out if you have a few minutes. Never seen one quite like it lol
https://t.co/fH4a9vPMUO
Interesting @derektmuller post on law faculty using AI detection in hiring, article reviews, and tenure & promotion. Thanks for taking this on, Derek.
I urge law faculty to be very skeptical about the Pangram estimates of AI-generated text, especially numbers like "28% AI drafted."
Pangram asserts a very low false positive rate in its evaluations, such that only 1 in 10,000 times is text labeled as AI when it was not AI-generated. But read the fine print. My understanding is that this is based on testing with pre-ChatGPT (2022) text. See: https://t.co/Lron48GqQ8
Post 2022, text is cross-contaminated: a lot of text in the world is a mix of human- and AI-generated text. There is also data drift, as human writing styles mimic AI generated text and copy-editing. See: https://t.co/x8bnwuBC8j
In other words, large-scale testing of the accuracy of Pangram on 2026 documents would almost certainly reveal a much higher false positive rate. I'm not aware of such testing. And there has not been independent third-party testing to support Pangram's numbers. I mention Pangram both because Derek mentions Pangram and Pangram asserts that it is the most accurate of the AI-detection tools. There are many AI-detection tools that perform worse, creating even more risks of false positives and false negatives.
Pangram asserts a false negative rate of about 1 in 300, again, when testing on pre-ChatGPT (2022) documents. See: https://t.co/UnRLdRRyUJ. It is obviously better to err on the side of a false negative, but we don't know how the errors are distributed across different groups. For example, there is a body of literature finding that AI detectors are biased against non-native English writers. See: https://t.co/ZUkuacRtaF
Also consider that there are many "humanizer" tools available to avoid detection of AI-generated text. Someone can use AI prolifically and use a "humanizer" to greatly decrease the probability that their use of AI will be detected. Sophisticated AI users can likely avoid detection.
Also, consider the ways in which a person uses AI. Say that Researcher A leverages AI extensively for brainstorming, research, outlining, and generating drafts, and then rewrites the text in their own words. Researcher B manually does brainstorming, research, outlining, and drafting and only at the end of the process uses AI to help improve clarity and conciseness. Researcher A's use of AI is highly unlikely to be detected. Researcher B's use of AI to polish text is far more likely to result in text that is labeled as at least partially AI-generated. Which use of AI is more objectionable to the anti-AI crowd?
Law faculty and law schools should not be making high-impact hiring, publication, and tenure & promotion decisions based on AI-detection scores. There is insufficient evidence of the validity and reliability of AI-detection in this context, and there is evidence that AI detectors are biased.
Based on the evidence, we should be very concerned about the fairness of using AI-detection for student and scholarly writing.
In my opinion, we should not discourage the use of AI by scholars in any event. Used responsibly, AI can facilitate much more rigorous research and raise the bar for the expected quality of scholarship. There seems to be a concern that scholars are entering a prompt into a system, generating a law review article, and submitting it for publication. While AI systems have improved greatly, in most cases (nearly all?) a one-shot article will be weak upon close inspection. Used responsibly, AI tools can augment the brainstorming, research, outlining, drafting, and editing processes, with the human in the loop playing a very important role. I don't see the role of the human in the loop changing anytime soon, including because we will soon begin to recognize that some scholarship is within the capability of AI tools and we won't value it as much. We will value scholarship that is the product of an AI baseline plus what a human contributes to the process. What counts as scholarship will change. What scholars do in a world with AI--just like what lawyers, doctors, software developers, and other knowledge professionals do--is evolving, rapidly.
Whatever you think of AI use in scholarship, the use of AI detectors for high-impact decisions regarding hiring, article selection, and tenure & promotion without stronger evidence of their validity and reliability in this context is highly problematic. The risks of using AI-detection tools outweigh the benefits, and--as discussed in the WaPo article above--the risks of errors will increase with the passage of time as AI- and human-generated texts converge.
Many staggering results here. But this is a particularly amazing one: the exponent for matrix multiplication is no more than 2.25. The previous world record had been something like 2.37. This leap in progress is like Bob Beamon's long jump.
https://t.co/M4Ni3HvZf0
Quasi-RH?!?!???! Are you kidding me? If a human did this, it would be an instant Fields Medal, no questions asked.
RH says zeta has no zeros in Re(s)>1/2. The best we had until a second ago was a region that got thinner and thinner the higher up the imaginary axis you go. I thought maybe they’d fatten that up a bit, that’d be a massive breakthrough. But no. They got a zero free strip!!!! Insane
We’re releasing a broad range of new mathematical results produced by an internal frontier model.
We’ve been consulting with the independent Advisory Group on Mathematics and Artificial Intelligence at the Institute for Advanced Study, and we have drawn on their advice and public recommendations to inform how we release these results.
https://t.co/7N6TPlft1P
Introducing Kardashev-0.7, the world’s first trained swarm made of 32 distinct models
Trained with RL for Population Scaling (RLPS), 32 models organically develop specialization & complementary capabilities, delivering frontier performance at:
- 0.007x ~ 0.02x of the inference cost
- 0.03x of the required memory
Civilization advances through different minds specializing and working together. We’re bringing that principle into AI
@BanburyRoadAI, we’re scaling intelligence by model count, toward civilizations of models that learn to build on one another
To justify today's AI capex, we'll have to spend 9% of GDP on AI services, a study finds. Is it plausible we'll spend as much on AI as on food? Twice as much as on energy? The law of diminishing returns would like a word. My column: https://t.co/Vsi9AE3WUC
New research: we hired 12 CPAs to complete four realistic accounting tasks, then had AI models from the past few years complete the same work.
As of May 2025, accountants outperformed AI. Now they lose out to today's frontier models, which are near perfect on these tasks. 🧵
From the NFL: “The NFL appreciated the opportunity to meet with Governor Pritzker and legislative leaders today. Timely action is essential as the Bears work to find a long-term stadium solution for the club and their fans.”