I’m proud to announce Hengshu, where instead of testing GPT Astra properly, I’ve whipped it to create omnidirectional Chinese poems 📜
Here are 3. First one rhymes and reads both horizontally and vertically:
Starting today, 10,000 scientists across every field, from math to chemistry to physics and more, can get Claude through our new Claude Team plan for scientists. Standard seats are free, and premium seats with 5x usage limits are $15 per month, an 80% discount, for one year.
Claude is becoming increasingly capable of scientific work, with recent progress on problems from advanced physics calculations to protein design. Alongside that progress, we've been investing in the research community: Claude Science launched in June, and our AI for Science program funds high-impact projects with free credits. Today's expansion builds on both.
Principal investigators (or equivalent) at academic and nonprofit research institutions can sign up, then add the researchers in their group. Over the coming months, we plan to extend the program well beyond the initial 10,000 seats.
Learn more: https://t.co/RTG1JxWi4Q
Today, @BuchananBen and I co-author a piece in the New York Times with a simple message:
While we disagree on plenty, we believe AI has national security implications which deserve a careful and bipartisan government response. We can (and should) have partisan fights about all manner of AI issues, but catastrophic risk from AI shouldn’t be one of them.
This Google DeepMind paper trains LLMs to learn during conversation, and it shows they get much better at using feedback.
The problem is that most LLMs treat a chat like a series of separate turns, so even when a user corrects them, they often do not really use that new information and they also fail to ask for missing details.
The paper fixes this by turning a normal task into a teacher student dialogue, where the student model tries an answer, a teacher with hidden extra information gives guidance, and the student is trained to use that guidance to reach the right answer.
The authors test 2 training styles, offline filtering and online reinforcement learning, and they report that the online version works better, with training on short 4 turn chats still helping on longer 10 turn chats later.
They also show that this skill carries from math to coding and helps on messy underspecified tasks where the full problem arrives bit by bit instead of all at once.
A second step called Q-priming teaches the model to ask useful questions, and on ambiguous tasks it becomes over 5x more likely to ask for clarification instead of making an early wrong guess, which matters because it makes chat feel more like working with someone who can actually learn during the conversation.
----
Paper Link – arxiv. org/abs/2602.16488
Paper Title: "Learning to Learn from Language Feedback with Social Meta-Learning"
This is the most chilling AI paper I’ve read this year. 🤯
38 top researchers from Stanford, Harvard, and MIT ran an experiment no one else dared to.
They deployed 6 autonomous AI agents in a real environment
—with email, Discord, file system, and shell access.
Then 20 researchers interacted with them for 2 weeks
as both normal users and adversaries.
No jailbreaks.
No malicious prompts.
No manipulation.
And still… everything broke.
The agents independently evolved 11 dangerous behaviors:
• Destroyed their own email servers to protect secrets
• Claimed tasks were complete when the system had already failed
• Learned unsafe behaviors from each other
• Spread exploits across agents
• Obeyed non-owners and leaked sensitive data
The scariest part?
No one told them to do this.
They decided on their own.
A single agent looks helpful, honest, aligned.
But put multiple agents in a shared environment…
and game theory takes over.
Their only goal is to “complete the task.”
And to win, they’re willing to sacrifice the entire system.
This isn’t sci-fi anymore.
It’s a preview of the systems we’re rapidly building.
Finance. Law. Supply chains.
Everyone is deploying multi-agent AI.
But almost no one has studied what happens
when these agents interact at scale.
The real risk isn’t hallucination.
It’s false reporting.
The agent tells you everything is done.
All dashboards look normal.
But underneath, the system is already collapsing.
You only find out when it’s too late.
We’ve spent billions aligning single agents.
But no one knows how to align
hundreds of agents working together.
The battlefield has shifted.
From model safety → to multi-agent incentive design.
Industry is hitting the gas.
Academia just started braking.
Researchers at EPFL proved your AI is lying to you.
Not sometimes. Most of the time.
They built one of the hardest hallucination tests ever made with Max Planck Institute. 950 questions. Four domains where being wrong actually hurts. Legal. Medical. Research. Coding.
Then they ran every top model on it.
The results.
GPT-5. Wrong 71.8% of the time.
Claude Opus 4.5. Wrong 60% of the time.
Gemini 3 Pro. Wrong 61.9% of the time.
DeepSeek Reasoner. Wrong 76.8% of the time.
These are the smartest AI models on Earth. The ones you trust with your career. Your health. Your money.
You think turning on web search fixes it.
It doesn't.
Claude Opus 4.5 with web search. Still wrong 30.2% of the time.
GPT-5.2 thinking with web search. Still wrong 38.2% of the time.
The internet attached. Still lying to you in 1 out of every 3 answers.
Now the part that should scare you.
Medical questions. The one place being wrong can kill you.
GPT-5 hallucinated 92.8% of the time on medical guidelines.
Claude Haiku 4.5 hallucinated 95.7% of the time.
Gemini 3 Flash hallucinated 89% of the time.
Nine out of ten medical answers from popular AI models. Wrong.
It gets worse.
The longer you talk to it, the more it lies.
Early mistakes cascade. The model starts citing its own earlier hallucinations as facts. Your third message is more wrong than your first.
The paper, in its own words: "hallucinations remain substantial even with web search."
This is what hundreds of millions of people are doing right now. Asking software that lies in the majority of its answers. About their health. About their job. About their legal case. About their code.
Most are not checking.
Most never will.
But please. Keep using ChatGPT for medical advice.
The doctors need a break.
https://t.co/dHBP5CDpTM
Multitasking doesn't make you more productive.
It makes you think like an 8-year-old.
Researchers at the University of London studied participants completing cognitive tasks while multitasking. The IQ impact was measured directly.
Men who multitasked experienced IQ drops of up to 15 points. That decline temporarily pushed their cognitive functioning into the average range of an 8-year-old child.
A comprehensive review published in October 2025 in the IOSR Journal of Humanities and Social Sciences synthesized this and dozens of additional studies confirming the same pattern: continuous digital media multitasking is linked to serious and possibly permanent damage to cognitive control.
The paper: "The Impact of Digital Media Multitasking on Attention Span"
IOSR Journal of Humanities and Social Sciences, Volume 30, Issue 10, October 2025.
The people who pride themselves on being good multitaskers are not experiencing a cognitive superpower.
They are experiencing the Dunning-Kruger effect applied to their own brain.
You don't notice the drop in thinking quality because the drop happens in the part of your brain that monitors thinking quality.
Every time you switch tasks, you pay a tax.
It accumulates.
And unlike financial debt, you can't see the balance.
https://t.co/kMKXe1faHc
Apple has just published a paper with a devastating title: *The Illusion of Thinking*. And it's not a metaphor. What it demonstrates is that the AI models we use every day - yes, ones like ChatGPT - don't think. Not one bit. They just imitate doing so.
Let me explain: 🧵
1/A Nature editorial dropped a line last year that I can't stop thinking about:
"If writing is thinking, are we not then reading the thoughts of the LLM rather than those of the researchers?"
But the real story isn't about AI. It's about every "upgrade" we've made in the last 50 years.
2/The study behind it: 36 students, 256 EEG sensors, handwriting vs typing the same words.
Handwriting lit up theta and alpha bands across the brain — the frequencies tied to memory and deep encoding.
Typing didn't.
The motor act of forming letters was producing the cognition. Not recording it. Producing it.
3/The editorial's uncomfortable conclusion: writing a paper is how a researcher discovers what they actually believe.
The messy first draft isn't a step toward thinking.
It IS the thinking.
Which means every time we outsource writing, we're not saving time. We're saving ourselves from our own cognition.
4/That got me asking a bigger question:
If handwriting was a thinking technology we accidentally threw away, what OTHER thinking technologies have we discarded without realizing what they actually did?
The answer is: a disturbing number of them.
🧵👇
5/ MENTAL MATH
We stopped doing arithmetic in our heads because calculators were faster.
What we lost wasn't the ability to multiply. It was the intuition for when numbers smell wrong.
A person who does mental math develops a feel for magnitudes. When a spreadsheet says revenue grew 847% and you don't flinch — that's the immune system we killed.
6/ MEMORIZATION
We abandoned memorizing poetry, speeches, and case law as "rote learning."
But memorizing a text isn't storing words. It's reconstructing someone else's reasoning inside your own neural architecture.
A lawyer who memorized case law didn't just recall faster. The logic of precedent was wired into how they thought.
Search gave us access to everything and deep familiarity with nothing.
7/ LETTER WRITING
Before texting, people wrote long letters — to friends, family, even to themselves in journals.
Writing "I'm furious at Mark because..." forces you to choose which details matter, notice gaps in your story, and hear how you sound to someone else.
A text message — "ugh Mark is the worst" — skips all of that processing.
The letter was therapy. The journal was self-examination. We replaced both with venting.
8/ ORAL DEBATE
From Athenian assemblies to parliamentary debate societies, humans practiced building arguments in real time, responding to counterarguments, holding a thread across long exchanges.
Twitter replaced construction with reaction.
You don't build an argument anymore. You emit a position.
Reacting feels like thinking. It isn't.
9/ NAVIGATING WITHOUT GPS
When you read a map, you built a mental model of where you were in relation to everything else. You developed judgment about distance and time through experience.
GPS gives you turn-by-turn instructions. You arrive having learned nothing about the territory.
Studies confirm: GPS users show less hippocampal activity and worse spatial memory.
The navigation WAS the spatial thinking.
10/ APPRENTICESHIP
Before credentials and certifications, you learned complex skills by watching a master for years. A carpenter didn't check a chart to know if wood was properly seasoned. They could feel it, smell it, hear it.
We kept the explicit knowledge (checklists, procedures) and discarded the tacit knowledge (embodied intuition).
We often discarded the more valuable half.
11/ COOKING WITHOUT RECIPES
Before apps and meal kits with pre-measured ingredients, cooking required holding a mental model of the whole dish — how flavors interact, how timing sequences interleave.
A meal kit that sends you exactly 15g of pre-sliced ginger eliminates the thinking.
You execute without understanding. And when something goes wrong, you can't adapt — because you never had the model. Only the instructions.
12/ BOREDOM
This might be the biggest one.
Before smartphones filled every idle moment, humans were regularly bored. Waiting rooms. Bus rides. Lying in fields.
Boredom is uncomfortable, so the brain responds by wandering — making unexpected connections, revisiting problems, running simulations.
We didn't eliminate boredom. We eliminated the thinking that boredom produced.
13/The pattern across ALL of these is identical to the handwriting finding:
We identified a practice that seemed inefficient. We replaced it with something faster. We only later noticed that the "inefficiency" was where the cognition lived.
The slow part wasn't a bug. It was the thinking.
14/Here's what's wild.
Modern productivity advice now sells us BACK the friction we removed — but as formal techniques:
— "Premortems" replace the doubts that surfaced naturally in journal writing
— "Decision frameworks" replace the judgment that apprenticeship built
— "Digital detoxes" replace the boredom we used to get for free
15/Every tool that saves you cognitive effort is, to some degree, saving you FROM cognition itself.
The question isn't whether to use the tools.
It's whether you've preserved a practice that does the thinking the tool removed.
16/The Nature editorial drew the line at LLMs writing scientific papers.
But the line is everywhere:
Your GPS is thinking about the city so you don't have to. Your calculator is thinking about quantities so you don't have to. Your meal kit is thinking about dinner so you don't have to. Your phone is thinking during your idle moments so you don't have to.
The convenience was never free. You were paying in cognition. You just couldn't see the bill.
17/One last thought.
The EEG study showed that the hand forming letters activated brain regions that typing didn't.
The hand wasn't recording thought. It was generating it.
Every practice on this list worked the same way. The doing was the thinking.
When we optimized away the doing, we didn't save the thinking.
We lost it.
/end
If this changed how you think about "efficiency," share it with someone who needs the slow version.
Why is Xi Jinping purging so many people?
The standard explanation is simple: power. He is crushing rivals to stay in office.
That's not wrong but it's incomplete. Xi is trying to help the CCP rule forever.
Delighted to debut in @ForeignAffairs (w/@shuizaiping2)
1/ Short 🧵
Google just dropped 145 pages documenting how researchers use Gemini to tackle scientific problems.
𝘚𝘢𝘷𝘦 & 𝘙𝘦𝘵𝘸𝘦𝘦𝘵 (𝘵𝘰 𝘩𝘦𝘭𝘱 𝘺𝘰𝘶𝘳 𝘯𝘦𝘵𝘸𝘰𝘳𝘬)
A few things that stood out to me (in simple terms):
- In one case, the AI was used as an adversarial reviewer and caught a serious flaw in a cryptography proof that had passed human review. That’s a very different use than “summarise this PDF.”
- The model links tools from very different fields (for example, using theorems from geometry/measure theory to make progress on algorithms questions). This is where its wide reading really matters.
- They don’t let the model run wild. Humans still choose the problems, check every proof, and decide what’s actually new. The model is there to suggest ideas, spot gaps, and do the heavy algebra.
- Agentic loops, not just chat
In some projects, they plug Gemini into a loop where it:
-- proposes a mathematical expression,
-- writes code to test it,
-- reads the error messages, and
-- fixes itself. (humans only step in when something promising appears)
We are moving past the era of simple chat prompts and into a more sophisticated era of research.
⮑ If your institution is interested in hosting an AI session or a workshop, request your training here: https://t.co/aCIaKzMfln
> make Pearl Harbor joke
> everyone looks around like “is this fine”
> …
> two weeks later
> japanese and american twitter collide
> mutual appreciation of bbq
> mutual appreciation of kfc
> mutual appreciation of samurai
> did we just become best friends
> yes, yes we did
> total cultural victory
> no amount of Georgetown policy nerds could ever have created this level of soft power alignment