Such feats are coming to the humanities & social sciences too - esp. the pre-print era, with its mostly destroyed, scattered & difficult materials...
Find out more, or showcase yours: Ars Inquirendi, 20-22 Nov @StEdmundHall & Online!
CfP closes 31 Aug. https://t.co/EObqmnxNgj
An internal version of Astra, @OpenAI’s next major model family, solved 10 major open problems in mathematics, quantum complexity, and theoretical computer science.
We believe it will be a major step for scientific reasoning. https://t.co/iP6cyheZ7i
@blagden_david The entire History faculty of the University of Hertfordshire just got fired. And this will continue for as long as the sector is unable to gain political support. And this stuff doesn't help, at all.
The study author @stefanski_radek will be speaking at Ars Inquirendi, 20-22 Nov @StEdmundHall & online, to tell us all about it ! Registration opens soon
New study using AI to analyze content of thousands of texts over time (a method which could revolutionize intellectual history, and which I suspect we'll see more of) shows that that between 1000 and 1920, cultural resistance to new ideas fell by half, which appears to have given a huge boost to economic productivity. By Radoslaw Stefanski: "Cultural Capital and the Productivity of Ideas: Evidence from Historical Texts." https://t.co/EhZOSih85y
An analysis of 3,500 couches featured in IKEA catalogues from 1960 to 2021 shows that European interiors became noticeably greyer in the 2010s. Design trends of course don't last forever. IKEA is now bringing long-forgotten colours and patterns back from the archives. A more colourful future awaits your living room!
The middle L in LLM, the economic miracle of our era, stands for Language. So naturally @UniofExeter's new redundancy list is 85% humanities - 445 staff at risk, medieval studies included. Tell them not to commit economic suicide. Sign here: https://t.co/bVQLCBdLgG @UKChange
Several newly appointed British Academy fellows at the University of Exeter have been told they are at risk of redundancy, just weeks after learning they had achieved the honour #jobcuts#redundancies#highered
https://t.co/6SxtkJge72
@arsinquirendi Both, depending what you ask. As a census of belief, the pre-print record distorts - thin, elite, curated. As a record of what each generation kept and taught, it's the object I'm measuring. The results don't hang on it- rebuilt with only post-1500 works, the numbers barely move.
An interesting corpus analysis, with limitations: looking at history through the lens of printed books, does it aid or distort our understanding of the West (or any culture) before the age of mass print?
Everyone from Weber to Mokyr puts culture at the center of the rise of the West. But nobody has had a long-run cultural series you could drop into a growth model.
So I built one: LLMs read 23,000 books from the Western canon, scoring what each endorses.
Year 0–1920, one chart:
‼️ Anthropic was hiding a secret internal project to destructively scan every book in the world, irreversibly damaging heritage. The company spent billions of dollars to slice the spines off millions of books, even very rare ones. This means no other AI, and no human, can access this knowledge anymore.
It now markets itself as the responsible lab.
One bookseller told 404 Media his inventory is full of rare, foreign-language and low-circulation books, meaning that if they are destroyed in the process of becoming training data, they will be even harder to obtain.
Second-hand booksellers across Europe are sounding the alarm, because rare books are being destroyed. They have received random lists of thousands of titles. Books too obscure to resell at a profit. Requests arriving overnight from companies nobody in the trade has heard of.
One email came from "Nataly" at 2077AI, a Singapore company. Attached: 3,000 English titles grouped by ISBN. A 1999 geomechanics monograph. A 2018 study of laser shock peening on ceramics. A 2021 academic book on Irish folk tales.
A German dealer watched the orders arrive every night between 3am and 5am. Canadian company Zoom Books, buying systematically, titles with nothing in common. Zoom told Swiss broadcaster SRF that this is a regular recycling and trading model.
We only know about this at all because of litigation. The brokers sourcing the print books for AI training advertise NDAs on every engagement; buyer identity is never disclosed. So the real question is: how many labs are doing this that no lawsuit has yet exposed?
Buying books by ISBN tells you about the title, never the copy. So no one can know if the volume guillotined for AI training is inscribed or annotated - in other words, unique and gone. Has @AnthropicAI clarified whether any copy-level check happens at all?
🦔AI companies are bulk-buying rare books, scanning them through high-speed machines that cut the spines off, and shredding the originals. A service called ISBNdb facilitates orders of up to a million books and keeps buyers anonymous. Pre-2022 books are premium because they're free of AI-generated text. A federal judge ruled the practice is fair use because eliminating the original means only one copy exists at a time. Anthropic hired the former head of Google Books partnerships to obtain "all the books in the world."
My Take
This got to me. A bookseller told 404 Media that rare books with almost no surviving copies are being fed into this pipeline. Books that survived wars, fires, and centuries of handling are being shredded so an AI can learn to write a better marketing email.
ISBNdb's website literally says "'AI company destroys two million books' is not a headline that generates sympathy," and they still built an entire business around making it happen quietly. They offer NDAs as a feature. They coach clients to call it "digital preservation."
I've covered AI companies scraping the internet, torrenting libraries, and stealing music. This is worse because it's irreversible. You can re-upload a website. You can reprint a bestseller. You can't replace the last three copies of an 18th-century botanical text once someone shreds them for training data. And the judge said it's legal. So it's going to accelerate.
"We shred rare books and offer NDAs so nobody finds out" is a legitimate business model in 2026. What a timeline.
Hedgie🤗
Hello everyone. Please go show your support to the guys over here who have translated in full St. Bonaventure's Commentary on the Sentences, volumes I->IV. St. Bonaventure's major systematic work is now finally available in English. https://t.co/cXTOl6TDX3
The LLM-enabled blossoming of these new forms of edition is fascinating. I wonder whether they'll do to commercial critical editions what Wikipedia did to Britannica ?
Migne, Patrologiae Latina et Græca
❦
Every column of the Latin and Greek Fathers, readable and citable — and, work by work, for the first time in English
Pretty amazing this is online and free
https://t.co/iRoyBi1EHS
Dear oh dear. Fable's security safeguards decided to stop my overnight processing of 0.3mn centuries-old Persian manuscript records. Why? Probably because one of dozens of sub-prompts (themselves Fable-generated) included the dread word 'enrichment'. The filter is too crude!