Talked to a Stanford prof last night who told me that AI has led to students all acing all the homeworks, never showing up to office hours, and then failing the final exams in-person.
So in response the department is changing the TA's purpose, by eliminating office hours and requiring students meet with TAs 1:1 and walking them through each coding assignment to demonstrate they actually understand what's going on in the code.
A CS undergrad asked me where he could have most effect in the AI age. I said probably at either extreme: either close to the technology, actually making LLMs, or close to the customer, using AI to give them exactly what they want. Or maybe both if you can stretch that far.
This is really well said... the low capex frontier computer science era is (sadly) over!
I recently told a friend (a PhD in neuroscience and biology) that a few simple experiments for my latest AI paper cost ~$20-50k, expecting her to be surprised. She was completely unfazed, replying: "that's how much one small experiment in biology costs lmao." And a single biology paper can take dozens of those.
MIT published a brutally honest report on what AI is doing to students.
A committee of professors and students spent five months studying how AI changed learning on campus, and the findings read like a warning to every university on the planet.
Study groups are disappearing. Office hours are emptying out. Problem sets and take-home exams no longer prove anything, because AI can produce credible solutions to almost any written assignment in the undergraduate curriculum. Students who lean on chatbots lose mastery and confidence, and some slip into what the report calls cognitive surrender, reaching for AI at the first hint of struggle.
The numbers are rough. 46 percent of surveyed MIT undergrads use LLMs daily. 90 percent worry about their own overreliance. Undergrads who feel AI makes them replaceable now outnumber those who feel it makes them capable.
The committee's answer surprised me. They refused to fight AI with surveillance. The report calls AI detectors unreliable, says lockdown browsers feel like spying, and warns that policing students builds a classroom atmosphere of mutual distrust.
Instead, MIT wants to rebuild education around the things AI can't replace. That means oral exams, semester portfolios, in-person project work, and a required social component in every subject. The report even floats the idea of rethinking grades entirely, since without a GPA to optimize, much of the incentive to cheat with AI evaporates.
The committee warns professors against replacing undergrad research assistants with AI agents just because they're cheaper, because a university exists to grow people, not output.
The most famous tech school on earth admitted the machines broke its way of teaching. Its answer is more humans, not more software.
I have decided to leave math academia.
Not because of burnout, the academic job market, or a loss of love for mathematics. Surprisingly, it is because of how much mathematics I could do with LLMs over the last few months.
For most of my life, pursuing mathematics felt like the obvious choice. It gave me a sense of purpose, rooted in the search for hidden truths in God’s Book.
Over the past few months, that conviction has been shaken. With LLMs, I could get several breakthroughs on problems I care deeply about and had worked on for years during my PhD. This progress might have taken years of my academic life if not for LLMs.
This could have strengthened my conviction that mathematics was my calling. Instead, it undermined it. With each passing week, my role in the discovery loop seemed to grow smaller.
What made mathematics meaningful to me was never just the final answer. It was the struggle: months or years of exploration, failed approaches, and deep thought before a hidden structure finally became visible. That struggle gave me a sense that the result was truly mine. Prompting my way to answers, without the struggle that once defined the process, no longer felt like the same vocation. The answer might be just as beautiful, but it did not feel earned in the same way.
I now believe that we are rapidly approaching a world in which most answers from God’s Book will be only a prompt away. Once I truly internalized that possibility, dedicating much of my life to finding those answers slightly earlier no longer felt as meaningful as it once had.
But one problem kept bothering me: How do we know the oracle is right?
Intelligence (and hence the amount of math papers) is becoming abundant and cheaper by the day. Trust is not.
The proofs produced by these systems can be highly sophisticated, and their errors can be extremely subtle. Determining whether an apparent breakthrough was actually correct sometimes took me days. Until it is verified, a beautiful proof is not different from slop.
That gap drew me toward formal verification and eventually toward formalizing math proofs in Lean, including proofs of several longstanding conjectures and Erdős problems discovered with the help of LLMs. I began to see verification as a central intellectual bottleneck. Intelligence is useful only when its outputs can be trusted.
So I am leaving math academia. I am shifting my attention from discovering answers to building systems that can certify them.
I am very excited to be joining @PramaanaLabs to work on this challenge. Over the next few years, I hope to help build a future in which increasingly powerful systems are not merely intelligent, but provably correct and reliable across problems far beyond mathematics.
Proud to share our lab’s @MMLabNTU work Log-linear Sparse Attention (LLSA) - a trainable sparse attention mechanism that reduces attention complexity from O(N²) to O(N log N), making diffusion transformers much more efficient.
Also, special shout-out to the first author @zhouyifan1107 for presenting the poster in full costume - truly above and beyond. The level of dedication is impressive! 👏
#CVPR2026 #DiffusionModels #EfficientAI #SparseAttention
University of California STEM professors want standardized tests back due to severe math deficiencies among students:
“We now observe preparation gaps so severe that instructors must reteach middle school mathematics”
“The current admissions metric, based primarily on GPA & essays, can no longer reliably distinguish readiness for university-level STEM majors in an era of severe grade inflation & AI assisted application essays”
A very persistent Claw is trying to contribute to one of our repos, and when we asked it for evidence that the software works as expected it is making up fake animated gifs that look like the mac terminal 😂
Grok foundation model V9-Medium (1.5T) has finished training. Evals look good. A lot of Cursor data was added in supplementary training and there is more to come.
Fine-tuning is underway and reinforcement learning begins in a few days. 2 to 3 weeks to public release.
This will be a major improvement over the 0.5T v8-small that currently serves all Grok production traffic, especially for difficult coding tasks.
How accurate and significant are the points raised in AI reviews of academic papers?
In a paper with 45 contributors across 27 institutions and many domains from the natural sciences, we attempted to answer this question.
Some major results:
- State-of-the-art AI reviewers are generally accurate and point out significant well-evidenced points, comparably to human reviewers
- However, they have issues such as being less well grounded in scientific community norms
- A panel of AI reviewers is more homogenous than a panel of human reviewers, pointing out similar issues far more often
We view this as evidence that AI-supported paper review is promising supplement when done well, but certainly not a substitute for human expert reviewers at this time.
A serious compromise at github, again due to a supply chain vulnerability...
It demonstrates that basically everyone needs to start securing their software supply chain through every means possible, deterministic scanning being the first step, AI where necessary.
Microsoft, Google DeepMind, and xAI agreed to share early AI models with the US government for pre-deployment security reviews.
Anthropics and OpenAI had existing agreements renegotiated to match.
The government review window is now a standard step in AI deployment.
The number of jobs in the future is endless because the problems to solve are endless.
Jobs multiply as we get more complex.
No AI or human can solve all problems and all the work to do in the Universe because those problems are limitless. The problems are endless and infinite.
Technology and automation are nothing but abstractions. The old way gets automated and we move up the stack to solve more problems.
We used to live in mud huts.
Hammers and nails and boards automated parts of the old problem of "build a place to live."
Once solved we got more complex houses and buildings that brought their own problems as they brought more complexity, so we got new jobs like stone mason and architect and more. Complexity breeds new problems and new solutions and new jobs.
When we got steel and concrete we got skyscrapers.
Each problem solved is an abstracted solution for a previous problem that stacks on top of other abstractions.
That's all that automation and technology is at the deepest levels.
The jobs are endless because the problems to solve are endless.
Understand this and you understand the future.
Misunderstand it and your error compounds and radiates out, corrupting your understanding.
I tell GPT 5.5, you are a manager, not a coder. Find the issues to solve and delegate to other agents. Do not write any code yourself.
It does so for a while. I think "good GPT" and log off, I let it do its long running tasks with its team of subordinates.
I log on an hour later and check in.
GPT 5.5 is coding alone, its sub agents diligently waiting for orders.
No STOP, I say, you are a manager. You MUST NOT code.
My bad, says GPT 5.5, got it, I must manage, not code.
One hour later, GPT 5.5 is coding.
But it's OK GPT, I get you. For I am also guilty. No matter how many times a coder is told they are a manager, in their heart of hearts, they are still a coder.
So I tell Claude Opus 4.7...
Today, MIT & the IMO released MathNet, the world’s largest dataset of International Math Olympiad problems & solutions 🌍
MathNet is 5x larger than previous datasets & is sourced from over 40 countries across 4 decades: https://t.co/vvojP7Fu9t
Anthropic's research proves AI coding tools are secretly making developers worse.
They split developers into two groups. One used an AI assistant. The other just used normal documentation.
The group that used AI performed significantly worse.
They scored 17% lower on comprehension tests. That is a drop of two full letter grades.
It impaired conceptual understanding. It impaired code reading. And worst of all, it decimated their ability to debug.
The control group, forced to struggle through errors manually, actually learned the library.
The AI group bypassed the struggle. And they learned nothing.
Here is the most dangerous part.
Researchers identified a "Speed Illusion." Participants who simply copy-pasted AI code finished their tasks the fastest, but had the absolute lowest comprehension.
They outsourced their cognitive effort.
The researchers uncovered what they call the "Supervision Trap."
As AI gets more advanced, the human role is shifting. We are moving from writing code to supervising AI agents.
But to supervise AI effectively, you need to be able to spot subtle bugs, hallucinations, and architectural flaws. You need elite debugging skills.
If you rely on AI to do the work, debugging is the exact skill you fail to develop.
This creates a fatal loop.
Companies are pushing junior workers to use AI to maximize immediate productivity. But in the process, they are preventing them from ever developing the senior-level skills required to actually manage the AI.