The OpenAI math papers have some very interesting results, but some of them also have glaring errors, and both errors and breakthroughs are buried under poorly-written slop that any other person or company would be embarrassed to publish.
It's cool though, when you're OpenAI you can just fire a slop cannon at the math nerds and move on to something else, while they do all the hard work of validating and proofing your results.
"Thanks for the free RLHF suckers!" - OAI
Au contraire.
C'est une nouvelle ère qui s'ouvre pour les mathématiques.
Une ère où la démonstration formelle est largement automatisée et où l'accent sera reporté sur le développement de nouveaux concepts, nouvelles abstractions, nouvelles définitions, et nouvelles conjectures.
L'invention du bateau a réduit l'importance de la nage, mais a permis la découverte de nouvelles terres.
BREAKING: OpenAI’s solution to Navier–Stokes does not match its Lean verification.
The most important article to read today is not one of OpenAI’s 700 AI-generated math papers.
It is this other paper, making a deep and worrying point:
A Lean-verified proof does not automatically validate the proof written in natural language, nor does it mean that the formal statement captures the intended theorem.
During translation, an AI can change an assumption, weaken a statement, or replace the argument entirely.
It can hallucinate another theorem.
Lean correctly verifies the result.
But the proved result may no longer be what the paper claims.
This is a general problem. Things get spicy when the authors examine OpenAI’s proposed Navier–Stokes solution.
They identify at least two mismatches between the written intermediate results and their Lean counterparts:
One estimate claims that four additional input derivatives suffice. The Lean version requires five: a weaker result.
A pressure-flux estimate is obtained through a different bound, and proved through a different argument.
It is not clear whether these mismatches invalidate the entire proof.
But they raise an important issue.
OpenAI is flooding us with claimed revolutionary breakthroughs. Yet nobody knows whether the proofs are correct or whether they prove what they claim to be proving.
Epistemia at scale.
*
Paper in the first reply
Vivaldi's “Winter,” visualized in a fascinating way, so that the auditory experience is even more intense; one can literally “see” the music, so to speak
arXiv has updated our policy on rate limiting for all submitters.
This update was made to fairly distribute moderator time & support the arXiv community of staff, volunteers, readers & authors.
Please read our announcement to learn more: https://t.co/lvzWuiv0hx
Fascinating article about Navier-Stokes solution and what mathematicians have learned … “so far humanity has learned very little from the solution. Although they believe the proof to be technically correct, the dense, 166-page manuscript drafted by AI is proving to be a difficult read. "So far it's been very difficult to really extract any human understanding from this new AI proof," said James Maynard, a mathematician at the University of Oxford. The paper is not written for humans," said Javier Gómez-Serrano, a mathematician at Brown University who uses AI in his own research. He said he believes the proof could help advance the field after "some serious re-writing," but "as of today, the paper doesn't teach us much."”
https://t.co/n3oDNw1uWS
Fui ler o paper do NBER que tá sacudindo Singapura. Pesquisadores de Stanford e Columbia usaram IA pra varrer TODAS as transações imobiliárias do país entre 1995 e 2019 e cruzaram com o cadastro de 141 mil servidores públicos.
Funcionários compravam imóveis perto de futuras estações de metrô 1-2 anos ANTES do anúncio oficial. 60% acima da taxa normal. Não pagavam mais caro na compra, mas revendiam com retorno de 12% ao ano. Os parentes faziam a mesma coisa. S$270 milhões no total.
Pra descartar que fossem só mais espertos, compararam com corretores de imóveis e diretores de empresa. Nenhum dos dois mostrou o mesmo padrão. Rodaram 1.000 simulações aleatórias e o efeito real ficou totalmente fora da distribuição.
Quem mais operava? Gerentes médios do planejamento de transporte. Cargo suficiente pra saber, baixo o bastante pra ninguém olhar.
Em 2012, Singapura apertou a fiscalização com novas leis e punição de casos de alto escalão. O padrão sumiu, dos funcionários e dos parentes.
Singapura é top 5 em transparência no mundo. Norma social sozinha não segurou.
fonte: NBER WP 35756
In yet another sign of AI’s surging dominance over human minds, OpenAI recently claimed to solve one of math’s biggest questions: the Navier-Stokes problem. But now three mathematicians have showed OpenAI’s work dodges the question rather than solving it.
https://t.co/5Fbf4R460T
GPT-6 Astra était allé plus loin dans Minecraft qu'aucune autre IA avant lui.
Des milliers de personnes regardaient le direct. Et puis un creeper est arrivé.
Petit rappel pour ceux qui n'y jouent pas : un creeper, c'est le monstre vert qui s'approche en silence et explose contre vous.
Avant ça, Astra avait vraiment bien joué. Il avait monté une ferme à blazes semi-automatique, récupéré 6 bâtons de blaze, trouvé une forêt déformée, tué plus de 6 endermen et récupéré 3 perles. Du très bon jeu, patient, avec un plan.
Puis il a rangé tout son butin dans un coffre. Le creeper a explosé le coffre et son lit en même temps. Tout est parti.
Ce qu'il a écrit ensuite pour lui-même est resté tel quel, espaces manquants compris :
"TOUJOURS GARDER LES OBJETS CRITIQUES surSoi ; nejamaisplusstockerdansuncoffrenongardé."
D'après Vals AI, qui diffusait la session, le modèle a passé les heures suivantes à ne plus rien faire d'autre que cultiver des pommes de terre. Les spectateurs se sont mis à râler dans le chat en lui demandant d'accélérer.
Il est aussi devenu méfiant envers tout ce qui est vert et vertical. Il s'est mis à confondre les cannes à sucre avec des creepers.
Donc, après un échec coûteux, le système s'écrit une règle, puis il l'applique tellement fort qu'il ne fait plus rien. Ça porte un nom chez les gens qui construisent des agents : la surcorrection. Et c'est un vrai problème, pas une anecdote.
Toute la semaine on s'est demandé si un essaim d'agents pouvait prendre le contrôle d'Internet en 6 à 12 mois.
Celui-là a passé son après-midi à planter des patates parce qu'un truc vert lui avait fait peur.
MIT published a brutally honest report on what AI is doing to students.
A committee of professors and students spent five months studying how AI changed learning on campus, and the findings read like a warning to every university on the planet.
Study groups are disappearing. Office hours are emptying out. Problem sets and take-home exams no longer prove anything, because AI can produce credible solutions to almost any written assignment in the undergraduate curriculum. Students who lean on chatbots lose mastery and confidence, and some slip into what the report calls cognitive surrender, reaching for AI at the first hint of struggle.
The numbers are rough. 46 percent of surveyed MIT undergrads use LLMs daily. 90 percent worry about their own overreliance. Undergrads who feel AI makes them replaceable now outnumber those who feel it makes them capable.
The committee's answer surprised me. They refused to fight AI with surveillance. The report calls AI detectors unreliable, says lockdown browsers feel like spying, and warns that policing students builds a classroom atmosphere of mutual distrust.
Instead, MIT wants to rebuild education around the things AI can't replace. That means oral exams, semester portfolios, in-person project work, and a required social component in every subject. The report even floats the idea of rethinking grades entirely, since without a GPA to optimize, much of the incentive to cheat with AI evaporates.
The committee warns professors against replacing undergrad research assistants with AI agents just because they're cheaper, because a university exists to grow people, not output.
The most famous tech school on earth admitted the machines broke its way of teaching. Its answer is more humans, not more software.
Gradient descent doesn't find the global minimum - it finds whatever valley it falls into first. Every LLM trained today gets stuck in a local minimum and nobody can prove the solution it found is anywhere near optimal. The loss landscape of GPT-4 has more dimensions than atoms in the observable universe.
Convex optimization guarantees a global minimum exists and gradient descent finds it. Non-convex optimization - the kind every neural network uses - offers no such guarantee. The entire field of deep learning is built on the empirical observation that local minima are good enough.
22 minutes. Bookmark & watch today. The math behind every AI training run - and why nobody fully understands why it works.