New HuggingFace paper argues that increasing agent autonomy can gradually make human oversight ineffective by causing approval fatigue, overreliance, loss of situational awareness, and skill degradation.
As agents do more, users are pushed into approval mode: skimming plans, granting permissions, and reconstructing what happened across steps.
Over time, automation bias, approval fatigue, weaker situational awareness, and skill atrophy can make those approvals less reliable.
Worse, weak approvals can become training or evaluation signals, rewarding systems for being easy to approve rather than easy to scrutinize.
Their answer is cognitive scaffolding at 2 levels: developers add strategic friction, better approval design, behavioral monitoring, and checks that force attention at consequential moments.
– arxiv. org/abs/2608.23642
Title: "AI Agents Push Humans Out of the Loop"
LLMs are like Schrödinger’s cat: many possible trajectories, but you only see one outcome per run.
To really understand models, and debug where they go wrong, you can find the "forking tokens" that lead to different trajectories. Our new research does this 100x more efficiently!
New paper in Nature Human Behaviour makes two important points:
First, LLMs reduce the linguistic diversity of how people express themselves.
This is perhaps not super surprising at this stage. But the authors also show something deeper.
Since the words we use carry information about who we are, LLMs are also erasing the identity signals that we embed in our writing.
LLMs strip out the soul of our writing.
*
Full paper in the first reply
As I become a more senior PhD student, I feel like I’m getting dumber and my ideas less sharp. But on the bright side, I’ve gotten much better at expressing them with undeterred confidence. I guess my PhD training is preparing me well for a professor job.
Must-read article for anyone whose job involves *thinking*
https://t.co/88ezsrjMgz
"The problem with writing with A.I. is that it’s mentally enfeebling...
What, uniquely, does writing do? ... It compels thinking in ways that silent contemplation or spoken language rarely can."
Geoffrey Hinton says a big language model runs on about 1% of your brain's connections and still ends up knowing more than you:
"So in your brain, you have a hundred trillion connections, roughly speaking. Okay. That's a lot. And you only live for about two billion seconds. That's not much."
"If you compare how many seconds you live for, with how many connections you've got, you have a whole lot more connections than experiences."
"Now with these neural nets, it's sort of the other way round. They only have of the order of a trillion connections. So like 1% of your connections, even in a big language model, many of them fewer, but they get thousands of times more experience than you."
"So the big language models are solving the problem with not many connections, only a trillion. How do I make use of a huge amount of experience?"
"And back propagation is really, really good at packing huge amounts of knowledge into not many connections."
"But that's not the problem we're solving. We've got huge numbers of connections, not much experience. We need to sort of extract the most we can from each experience."
Two to three billion seconds is the whole budget. Everything you know, you learned inside it.
So evolution built you to squeeze a lot out of very little. Hinton's point is that a language model has the opposite problem and the opposite fix, and backprop turned out to be extremely good at that fix.
Worth noticing what this predicts about failure. A system running on 1% of your wiring and thousands of times your experience is not going to fail the way you do.
You fail from having seen too few examples. It fails from compressing too many into too little, and the compression is where the errors get made.
That is a strange thing to be deploying into hospitals and courts with no way to inspect it. We test these systems by asking them questions, which tells you what came out. Nobody can yet look at a trillion connections and say what got packed in.
- Geoffrey Hinton, Nobel laureate and Turing Award winner, on StarTalk (@StarTalkRadio) with Neil deGrasse Tyson.
Very insightful:
> The locus of judgment has flipped: users aren’t evaluating the machine (“Is it reliable?”) but themselves (“Am I more capable with it?”). That inversion explains adoption curves that reliability metrics never will.
https://t.co/JLuRKRgMkK
Trabalhos de auto-ML-research estão ganhando forma simultaneamente em diferentes empresas. Me pergunto se têm algum incentivo para serem usados como ferramentas robustas de experimentação, ou se tornarão ferramentas sem controle epistêmico algum dos seus usuários.
I work at Google DeepMind. This won't make me popular. But it's all public reporting:
2014: DeepMind reportedly sold to Google on conditions: no military use, independent oversight
2026: a Pentagon contract for "any lawful government purpose"
Not one safeguard survived intact
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying.
For months, I worked to stop this but watched powerful ethicists and institutions choose silence.
Here's what happened. 🧵
Today I'm publishing a new essay, Policy on the AI Exponential. AI is progressing extremely fast—much faster than the policy process was built to handle. The essay lays out where I think the technology is now, and the action needed to close the gap: https://t.co/Lh6PWae178
Bora plano A!
Existe uma falácia atualmente de que uma corrida armamentista de IA é inevitável.
Não precisa ser assim! Podemos caminhar rumo a um futuro mais aberto e distribuído, onde as pessoas comuns não são desempoderadas e não existe um único vencedor que leva tudo.
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power.
In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
Let's go Plan A! 我们选择 A 计划
There's a prevailing fallacy that the world must endure a bipolar AI arms race. It doesn't have to be this way!
We can push toward a more open, distributed future, one where the average person isn't disempowered and no single winner takes all.
In AI 2027, we predicted that AI would take over the world or irreversibly concentrate power.
In AI 2040: Plan A, we've laid out our positive vision for what should happen instead.
Concerned by AI 2027? Good news, the sequel is AI 2040! (In a highly optimistic future world)
I really like this type of futurism: try to rigorously forecast the future of AI, then flesh out a specific vision of this. This will be wrong in many ways, but useful to think about!
Chain of thought monitoring is one of our best safety techniques, and diffusion models might break it. But at least for DiffusionGemma, it turns out that we can recover most of the benefits! I would love to see similar transparency audits for any latent reasoning architecture