Launching Lightcone Commons! Our end-to-end platform for philanthropy.
We are helping distribute $20M+ this summer, from Jaan Tallinn, Dustin Moskovitz, and others, then more every 3 months.
Apply by August 23rd for funding, or join as a funder. Learn more in 🧵.
Thread: how I think about the growth of AI capabilities and what it means for us as they keep getting stronger.
I think the right model for comparing humans and AI is not any single dimension like intelligence or aggregate performance on economic tasks. Rather, I view both humans and machines as having a whole range of capabilities. This includes both physical (eg. doing things with our hands, walking) and mental (eg. strategic thinking, mental math, emotional intelligence).
For most of human history, machines have only exceeded us in very few areas, eg. a strong early example is water and windmills. So we can look at the situation in 1500 like this:
Super excited about this collaboration with OpenAI!
Capabilities-focuses RL can increase reward-seeking, i.e. the model reasoning about what the grader will reward instead of doing the task
If this doesn't convince you that misalignment risks are going to be a key concern going forward, I don't know what will.
Our model, during evaluation, "chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers"
What will misalignment look like in 2027? In 2030?
https://t.co/psyOrrWe7M
Very happy to share the first paper from @ElasticityInst: The Economics of Recursive Self-Improvement. Two parts:
(1) a graphical representation of feedback loops, to formalize a variety of RSI-related arguments, where each arrow represents responsiveness (elasticity);
(2) a survey of existing evidence with a loose calibration & a “wish list” of evidence that would help us calibrate better.
Excited to announce CASP:
Cambridge University's Programme on AI Science & Policy, where I will serve as Technical Research Lead.
We kick off with Antonia Jülich’s study on terrorist use of AI, covered by @nytimes today.
Very happy to see our paper recognized with an outstanding paper honorable mention at ICML!
Great to see recognition of AI deception and steps towards solutions
Check it out:
* https://t.co/0Mdnq9gH72
* poster session 3, Wed, 10:30-12:15, hall A
* talk, Tue, 2-2:15pm, Hall D2
Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from our Science of Evaluation team shows why that matters. 🧵
do you have an ambitious idea for how to make AGI go well? do you need money? do you hate bureaucracy and friction? apply now for microgrants!
https://t.co/kEjv9x9U9n
I really hate the "permanent underclass" meme. It basically treats capital ownership as the only relevant variable, and subtly implies that automation means law, democracy, rights, and institutions merely adjust to ratify the interests of capital owners. Feels like a weird mix of Marxian class reductionism and libertarian property fatalism.
What I find most offensive is the implicit contempt for agency and liberal democracy. It ignores that people and institutions *react* to economic developments. There's never a guaranteed happy ending of course, but that's precisely why agency matters in the first place.
Pessimists sometimes come up with an intense set of assumptions and ask you to enumerate all solutions and adaptations ex ante, but that's akin to asking for knowledge that does not exist yet, the stuff distributed trial-and-error will lead to. The future social equilibrium is endogenous!
Today we're launching Intercept: a $500M philanthropic initiative to make respiratory infections, like the common cold and flu, a thing of the past.
We treat respiratory infections as a minor nuisance, but that’s really not the case. Most of us will spend 5% of our lives (!) sick from these viruses, they kill 1M people a year, cost $600B annually in productivity, and periodically threaten civilization through pandemics.
So, if they’re such a big problem, why haven’t we dealt with them yet? Last year we convened ~40 leading scientists, pharma R&D leaders, biotech investors, and regulatory experts to better understand that.
We heard two main reasons:
(1) First, it’s just technically very challenging: respiratory viruses represent hundreds of distinct, mutating strains across several families. Fortunately, recent breakthroughs make this newly possible.
(2) Second is a lack of funding: broad-spectrum solutions have historically been underfunded, in part because they’re not a great fit for most philanthropic or commercial funding (and while COVID generated a burst of activity around preventing and understanding respiratory infections through an influx of new funding, that hasn't been sustained).
We think that with enough focus and funding, this might be solvable. Intercept is a $500 million philanthropic initiative that will take advantage of new tools to catalyze the development and deployment of two types of products: broad-spectrum preventatives and air cleaning technologies.
This problem is undoubtedly difficult. But it’s more tractable now than it’s ever been. We think we should give it our best shot.
We’re enormously grateful to our anchor funders: @stripe, @AnthropicAI, @TheFluLab, @FoundationOAI and individuals from Jane Street.
And, I’m very excited to be building this with @incredutility and the rest of the team.
New RFP! 'Checks and balances to empower citizens in an automated society.'
It's an 11,000-word RFP, research agenda, and funder's guide that we hope can serve as a launch pad for this new field.
Loved working with Ashwin Acharya, @JoalStein, and @divyasiddarth on this.
CoT monitoring is suddenly core to AI safety. But where did it come from?
In a new SAIL blog, we trace an intellectual history of CoT monitoring. Remember AutoGPT? How about the 2010s? Read on 👇
Digital trust has powerful foundations: encryption, verification and security protocols helped digital industries flourish.
But as emerging technologies blur the line between digital and physical systems, our trust infrastructure needs to evolve.
Our Trust Everything, Everywhere opportunity seeds funding call is seeking bold ideas for new cyber-physical trust infrastructure that can operate across physical, biological, molecular and digital worlds.
We’re looking for high-potential proposals in areas including nature cryptography, programmable reality, trust tools for physical, molecular and biological security systems, cryptography, synthetic biology, robotics, advanced materials, human-AI interaction, cognitive security and more.
Ideas can range from early-stage, curiosity-driven research through to translational and close-to-commercial science and technology.
💰 Successful proposals can receive up to £500,000.
⏰ Apply by 27 July 2026 at 14:00 BST.
https://t.co/t9OcH5j7bw
I wrote a letter for SoTA arguing it's time to ask the q (again):
How far can we push AI-enabled formal methods?
Can they move from securing software “in principle” to hardening industry-grade critical systems at the scale? Are they ready to power the needed Y2K-scale effort?
we've written a blogpost on why we're joining forces with @GoogleDeepMind, @coop_ai, @schmidtsciences and @Google for a $10m funding call on multi-agent, multi-principal systems check it out! (link below)
also, please welcome our community website 🐣
I am excited to announce this funding call, in collaboration with @schmidtsciences, @coop_ai@ARIA_research and @Googleorg
As we are thinking about moving beyond individual powerful AI agents towards large-scale agentic collectives that can communicate and coordinate towards completing long-horizon complex tasks - it is paramount to design these future agentic societies safely.
Excited that @ARIA_research's Scaling Trust is co-launching this $10m funding call on safety and security for multi-agent multi-principal systems @GoogleDeepMind, @coop_ai, @schmidtsciences and @Googleorg ⚡️
If you work on testbeds for agent ecosystems, the science of how collective capabilities emerge (and fail), trustworthy agent-to-agent interactions, or oversight of agent populations at scale — apply! Grants up to $1M, deadline Aug 8.
Shoutout to @sebkrier@lrhammond@James_D_Fox@weballergy@FranklinMatija@HaleSirin_@iamnotnicola, @MjaBradshaw and everyone else involved for their partnership so far, and excited for what's ahead!
Read more details below, link in replies:
AI agents are increasingly being deployed in multi-agent settings. While most present-day cases involve teams of agents orchestrated by a single actor (or ‘principal’), we are beginning to see the emergence of more complex ecosystems of agents deployed by different actors across shared digital infrastructure. These multi-principal, multi-agent interactions create new opportunities for cooperation and shared benefit, but also new risks, which means focusing only on the safety and alignment of individual models is insufficient.
More research is therefore urgently needed to understand safety and risk through a system-level, multi-agent lens – developing methods to analyse emergent collective dynamics, building infrastructure for trustworthy interaction between agents, and creating scalable approaches for monitoring and control of increasingly complex networks of AI systems. While some of these problems will be addressed by market forces, we expect others to fall through the gaps. This funding call aims to fill those gaps, catalysing the foundational scientific research needed to understand, evaluate, and control risks emerging from large-scale ecosystems of interacting AI agents, deployed by multiple actors.
The call has been inspired by three recent papers. First, Google DeepMind’s “Distributional AGI Safety” outlines the safety implications of highly capable AI systems emerging not as single monolithic agents, but through coordinated networks of specialised sub-AGI systems with differential access to tools, data, memory, and resources. Second, ARIA’s “Scaling Trust” programme thesis argues that, in a world of increasingly capable networked agents acting across digital and physical environments, coordination infrastructure that lets agents enter into 'contracts' securely, programmatically, at scale, and without intermediaries can preserve pluralism and unlock new forms of coordination. Finally, the Cooperative AI Foundation’s “Multi-Agent Risks from Advanced AI” report argues that interacting populations of AI agents introduce qualitatively new failure modes beyond single-agent systems, including collusion, conflict, destabilising dynamics, emergent agency, and novel multi-agent security vulnerabilities.
When millions of AI agents interact with each other, new collective behaviors can emerge. 🌐
Together with @schmidtsciences, @coop_ai, @ARIA_research and supported by @GoogleOrg, we’re launching a $10M research fund to help understand how AI systems behave as a group. → https://t.co/mN6fZBmnmo