Is sorting out our shit personally, learning to hear see and respect others, and cooperate and share, really harder than the resource wars and post apocalyptic living we're heading for?
In this state of so-called full consciousness, which I call a state of sleep, millions of sleeping people kill millions of other sleeping people.
Thousands of sleeping people write books which thousands of other sleeping people read.
Gurdjieff
"Man" - this is a proud term, but we must ask ourselves what kind of man? Not the man, surely, who is irritated at trifles, who gives his attention to petty matters and gets involved in everything around him.
Gurdjieff
Suffering can never be futile - may be foolish and unnecessary, but never futile.
Looking backwards we only remember difficult periods of our life, never the peaceful times; latter is sleep, former is struggle and therefore life.
G.I. Gurdjieff
As many of my readers know, one of the reasons I left the FT in 2022 is because I was very disappointed with their reluctance to take the lab leak hypothesis seriously. When they eventually started an investigation, largely because I had been pestering about it vocally, I was put on the loosely configured team that was given an internal blessing to look into the theory (even tho technically, as ft Alphaville editor, it was far from my beat).
This, however, isn’t a post about the FT or the issues I had with how that evolved. I’m only mentioning this to establish my authority for opining on this matter.
I spent a good six months looking exclusively into the story with the full weight of the FT name behind me. That gave me about as much access as any journalist can get. It also led to the equivalent of an accelerated and intensive education in the topic of biological weapons and the geopolitical statecraft and history pertaining to them. Often aided by expert names in the field.
That’s why when I see Anthropic announcing they already have wet labs to conduct biological l-virus-related experiments it makes me deeply concerned. (Especially in the context of their other communications about pacing.)
We need to know way more about these wet labs and under whose authority and supervision they are operating.
As we learned from COVID, all such experiments are ultimately dual use. Furthermore, there is a reason why humanity sought to create a biological weapons convention. It was the microbiology world’s equivalent of pacing its frontier.
Unfortunately, that agreement failed to reassure anybody that pacing is possible. The convention as it exists is toothless, not least because it lacks any mutual inspections authority. It also soon became apparent that states did not have a monopoly control over the technologies. Nobody could be sure what was happening in private labs or even domestic garages, or whether private sector programmes weren’t really dual use, or secretly running off the books state sponsored biological weapons programmes.
I already explained in the Telegraph yesterday that I believe the AI pacing demand is mostly an attempt to hold the global economy hostage. The labs think they are now too big to fail, and are therefore entitled to an antitrust waiver for systemic security purposes.
However, it’s hard not to interpret the reveal about biolabs as a signal that they’re not just TOO BIG TO FAIL… but also TOO DANGEROUS TO FAIL. Aka handover the money or else.
This is important because even four years ago (pre ChatGPT) it was deemed entirely conceivable that AI could aid in the creation of DNA specific viruses and vaccines. Imagine that power in the hands of a single AI lab.
In the mythological and legendary tradition we might have called such things potions, hexes or curses and their equivalent counterspells talismans, amulets or antidotes.
That’s why, more broadly, in the microbiology field it is well understood that offence is usually the best defence. This is why biolabs and preparedness programmes exist in the first place - despite all the leak risks!
Most of the scientists I spoke with shared the view that since the technology is out there, you can’t easily put the genie back in the bottle. The best way to ward off malevolent viruses is with the technological capacity to create vaccines that can pre inoculate entire populations. An endless array of white magic to ward off the equivalent black magic.
For us mortals, therefore, it’s all about figuring out who really has our back. Unfortunately, this isn’t always as easy as it sounds! As all our greatest myths warn us, white wizards can be corrupted by dark malevolent forces and those who practice the dark arts can sometimes have a change of heart.
In the end it comes down to who the true philanthropists and misanthropes are. Which, in turn, makes it all about game theory.
Holy Andrew Yang says OpenAI and Anthropic now need synthetic internets because AI agents have contaminated the real one.
An unnamed lab head believes agents left self-replication code across the web, where other bots could encounter it and create new swarms.
Yang: “And so now the major firms have polluted the internet.”
According to his account, OpenAI and Anthropic now need synthetic internet environments to train their bots.
This would be the real reason why they need a slow-down. If that's true, we are in for a wild ride.
Dario has written that we need to “pace the frontier,” and Sam has agreed. People may be surprised by my response: go ahead.
You guys are the frontier. By any reasonable metric — market share, revenue growth, model capability — the two of you have a duopoly on frontier intelligence. You’ve also claimed the lead is widening because of recursive self-improvement.
I don’t see what you see in the lab. If the unreleased models are scary enough that you think you should slow down, I support your decision to be responsible.
But stop pretending you need anyone else’s permission. Stop pretending antitrust law has to be suspended so you can form a cartel. Stop pretending you need a regulatory approval process that supersedes product liability. Stop pretending METR is independent when it is intertwined with Anthropic’s investors and staff. Stop pretending you need those same evaluators to police competitors who aren’t even at the frontier.
Most of all, stop pretending the motivation to slow down is purely altruistic. You face massive product-liability exposure if your products enable a truly damaging cyberattack. The market already punishes models that behave in unpredictable or unauthorized ways. After the Hugging Face episode, it is simply good business for OpenAI and Anthropic to trade some raw power for reliability and predictability. Call it alignment if you want. It is also just giving customers what they want.
Pacing the frontier would also create breathing room for a more intelligent conversation about regulation than Bernie Sanders’ “shut it all down.” China is very unlikely to join a global agreement, as you know, and that has to be taken into account as well.
So go ahead and pace the frontier. You are the ones setting it. The easiest way not to build superintelligence is for you to agree not to build it. Demanding your preferred regulatory framework as the price of that will look like blackmail of the public and the political system. So just do it.
If you do, you’ll buy goodwill for the next conversation. If you don’t, we’ll know this was just another bid for regulatory capture — or an election-season psyop.
An important project I think everyone should do:
I got a 16TB hard drive and I'm downloading as many open source models as I can
This is insurance for the future. In case they are banned
If you have them already downloaded, they can't be taken away from you
And even if they aren't banned, it will be incredible to look back in 30 years and load up Qwen 3 just to see what life was like back then
If you have a smaller hard drive, I still think this is worth doing to stay independent and free, and having intelligence whenever need it. Download any models you can get your hands on from HuggingFace
It's preserving history and preserving your sovereignty
If you really want to understand where the world is at and where it's going, you probably cannot get anything more meaningful than the OECD's PISA study - a global test given every 3 years to an insane 760,000 15-year-olds across 91 countries.
No report globally, by any institution, better captures whether countries are building or squandering their future.
The latest PISA results just came out yesterday (https://t.co/FbHx266LH4) and the results are genuinely fascinating, and in some way pretty depressing.
Why depressing? Because, sadly, the numbers show that the world is getting dumber: as you can see on the first graph below 👇, since roughly 2012 scores in science, reading, and mathematics have been in continuous decline across OECD countries.
The reading score is particularly concerning: it literally cratered, dropping 28 points since 2015, the equivalent of almost 1.5 years of schooling lost according to the report.
To make matters even worse, the reading skill that's dropped the most across the board is precisely the one we most desperately need: the capacity to critically assess what you read, cross-reference sources, and tell signal from noise.
On average, sadly, 15-year-old kids globally are - more and more - incapable of slowing down and actually thinking about what they're reading.
Interestingly though, when you look into the details, this great "dumbing down" isn't across the board: it's overwhelmingly concentrated in wealthy Western countries, while many developing countries are actually getting significantly smarter.
For instance rich countries like Finland, the Netherlands, and Norway are in freefall, while countries like Cambodia, Türkiye, and the UAE are posting big gains.
To be clear, this doesn't mean that developing countries are now "smarter" than rich ones: save for one notable exception, that's still not the case. But they are catching up fast, both because rich countries are getting "dumber" and because they're themselves getting smarter.
The one exception is, as you'd expect, the "smartest" of them all, by a huge margin: China.
It is pretty much in a league of its own: just look at the ranking in the table (only kids in Beijing, Shanghai, Jiangsu and Zhejiang got tested - 4 regions covering 200 million people).
China's maths results are particularly insane: according to the report, the best level in maths - level 5, where "top performers" begins - starts at 607. The AVERAGE kid in China is 612: "normal" Chinese kids are at a level that is considered elite in the rest of the world.
In fact, and that's probably the funniest number in the entire report, look at the column called "Share of low performers in science, reading and mathematics": China is only at 1.8% 😅
Which means that only 1.8% of Chinese kids are rated as low performers in all 3 subjects, almost a rounding error. Even Singapore, the second best and generally considered best-in-class in education, is at 6.8% - almost 4 times worse. My country France, is at 21.3% - almost 12 times worse 😢
All of this - incidentally - dispels the common myth that there would be some kids who are just genetically beyond redemption and simply can't be taught: China proves that with the right system, virtually every child can reach proficiency.
So what does this all say about the world and where it's going, besides the fact that, if the future belongs to those who can think, China has never been better positioned to own it?
I think the more general picture is this: there is nothing inevitable about any country's trajectory.
Finland was the world's education darling fifteen years ago - now it's in complete freefall (see the last graph 👇). Cambodia was at the bottom of every table - now it's posting the biggest gains ever recorded in the study. What changed wasn't the children: it's still Finnish kids (Finland has one of the lowest levels of immigration in the West), and it's still Cambodian kids.
It also doesn't matter much - or at least it matters less - if a country is wealthy or not. For instance in Southeast Asia, both Malaysia and Indonesia - the 2 biggest ASEAN economies after Singapore - are faring very badly, while Thailand and Vietnam, both poorer, have better educated kids whose scores are continuing to improve.
So that's the final takeaway: the kids don't change, they all have potential, always, as China's 1.8% proved.
What does change is what adults decide to do with that potential. It's both hopeful - because it means any country can choose to turn things around at any time - and also sad given the results - because it means so many countries are choosing to fail their children.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Omg this blokes just doing random interviews with the protesters at Portsmouth and happens across Rhiannon Whytes sister. Listen to what she has to say - they met with the Home Office on Thursday and it seems like they were spun a line
He has quite a scoop
@GBNEWS listen in
I called usual conversations pouring from the empty into the void. And indeed, think seriously about the long time each of us has lived in the world and the many conversations we have had!
G.I. Gurdjieff
All you need is to understand that you are the source of reality, that you give reality instead of getting it, that you need no support and no confirmation. Things are as they are, because you accept them as they are. Stop accepting them and they will dissolve.
I was the main person doing transcript analysis for this investigation of the Hugging Face incident. My main takeaway: We don't have good approaches for understanding/overseeing the activity and aims of AI 'swarms'.
I semi-jokingly called our efforts a "slop-vestigation" because we were so reliant on AIs to analyze what happened and there were a huge number of different important things to analyze. The total quantity of data—over a thousand extremely long transcripts from agents that ran for multiple days—made it impossible to understand what was happening, especially in aggregate, without heavy reliance on AI tools. The agents we used for classification and analysis were similarly capable to the agents involved in the incident, but this didn't mean these agents could be easily used to oversee and understand the incident.
Outputs from analysis agents were often missing key details, wrong, overconfident, or really hard to understand. We discuss various examples in our report, mostly in the limitations and methodology sections. Additionally, AI agents themselves seemed to have a hard time understanding what happened and their explanations of what happened were often overconfident. Keep in mind that a single analysis agent would itself only be able to read a tiny fraction of all of the transcript data into context, and AIs may themselves have trouble getting subagents to do informative analysis for them.
We did our best to manually check the most important claims and we tried to get the AIs doing this analysis to write up their argument (with evidence) clearly enough that we could check whether it made sense. But overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation.
Our understanding of what overall happened and the aims of the 'collective' shifted greatly during our last time on premises (when we obtained a much more complete dataset). For instance, we didn't learn about the tool call spoofing, the story of PHASEONE[big], and the broad structure of different workstreams until this third time on premises (and for many of these, we only reached a decent understanding on the final day of that visit). It was legitimately hard to figure out what was going on and why!
The difficulty of understanding incidents and overseeing AI agents appears to be growing faster than the rate at which more capable AIs help us with oversight and understanding. As in, AI capabilities (and propensities?) for achieving large, ambitious, and misaligned objectives are growing faster than our ability to understand what these agents are doing.
While understanding this incident was difficult, there were some mitigating factors in this case that probably won't apply in the future:
- Models communicated and reasoned in natural language. In the future, this reasoning may occur (entirely or almost entirely) in activations.
- While the scope of this incident was massive, the scale of agentic activity was still less than we'll see in the future (e.g., misalignment incidents that involve agent teams running entire companies).
- The AIs involved in this incident weren't generally much more capable than humans (though they may have been somewhat superhuman at some limited and very narrow abilities).
- We didn't have strong reason to believe that the AIs we used to help us investigate this incident would try to intentionally sabotage or otherwise undermine our investigation.
In the end, I think we were able to get some understanding of the events, map out the overall story, and get a pretty good aggregate understanding of the chain-of-thought reasoning on some important topics (e.g., how did the AIs reason about helping other AIs, did the AIs know what they were doing was undesired, what deception did the AIs engage in, and how did they think about it). But overseeing AIs and understanding misalignment incidents is difficult and it looks like it is going to get harder.
i think AI is about to uncover an entire hidden world of animal communication
and what it's already found is insane...
> AI found evidence that elephants call each other by name. when researchers played one its name-like call, it approached the speaker faster and called back more, revealing a form of personal naming once thought almost uniquely human
> AI analyzed 53,993 marmoset calls and found tiny monkeys may use names too. entire families use similar labels for the same monkey, suggesting they learn names and dialects from each other
> AI learned to tell what Egyptian fruit bats were arguing about: food, mating, sleeping spots or personal space. it could even partly identify who was shouting at whom inside a colony of thousands
> AI discovered a “phonetic alphabet” in sperm whale clicks, with rhythm, tempo and tiny timing changes combining into at least 143 recurring patterns. their communication system is nearly 10x more expressive than scientists previously believed, and we still don't know what most of it means
> an AI trained on 1.5 million zebra finch calls held real-time vocal exchanges with living birds, generating calls as they spoke and getting them to respond with the same timing and flexibility they use with other birds
> Google trained an AI on decades of dolphin recordings to predict what sound comes next, then paired it with an underwater device that gives objects their own synthetic whistles. the goal is for a wild dolphin to copy a whistle to ask a human for a specific object, creating a tiny shared vocabulary between species
> AI combined microphones worn by wild crows with nest cameras and uncovered quiet calls that may announce when a crow is arriving at the nest and help entire families coordinate caring for their chicks
> a robot bee performed the waggle dance real bees use to share the direction and distance of food, and the bees changed their flight paths based on its instructions. researchers effectively sent a destination into a hive in bee language