Thrilled to share that I am joining UC Berkeley as an Assistant Professor in the School of Information!
I start in Fall 2027, and I am recruiting PhD students this cycle. List me in your application if you're interested in frontier AI evaluation, AI policy, and AI's impacts on institutions such as science, law, and medicine.
I'm especially keen to work with students interested not just in high-quality research, but also in communicating it with a broad audience such as by public writing and policy impact. Fill out the form in the next tweet to indicate your interest.
As for this coming year, I'm moving to Berkeley this fall to start something new with @RishiBommasani and @random_walker. We'll have much more to share soon.
Hot take: I think it's still important to understand the code that our agents write!
In this mega thread (based on my AIE talk today), I will explain why that's the case, and show some ideas for how to efficiently understand code. Alright, let's dive in. 1/
Who gets access to frontier AI is now a geopolitical question. But another question matters too: which economies are most exposed to it?
Our new paper answers this by measuring national exposure to frontier AI across 141 countries.
(1/8)
from apps to material
software used to be something you opened
an app was a room with walls: calendar here, notes there, music there, work there. each one had its own logic, buttons, its own little kingdom. the user moved between kingdoms, carrying context in their head
but ai starts to break the walls
software becomes less like a destination and more like material. something you shape, combine, stretch, ask, remix, and leave behind as traces. a document can become an app. a conversation can become a workflow. a song can become a memory. a task can become an agent. the boundary between using and making gets blurry
the old model was: choose the right tool for each task
the new model is: express the shape of the thing you want, then refine it with the system you built
this changes the role of the interface. ui is no longer only fixed views for fixed functions. it becomes a surface where intent turns into structure. the best interfaces will feel less like menus and more like clay – responsive, persistent, inspectable, and alive
apps won’t disappear. rooms are still useful. but the deeper shift is that software stops being a set of sealed containers and becomes a medium people can think through
like paper, but executable
like language, but spatial
like memory, but programmable
software stops being something only programmers make
it becomes material anyone can shape
eagle is too clunky, mymind is too ugly, arena is too slow, freeform is too buggy, figjam is too web, cosmos is almost perfect, but not local. all i needed was apple photos app, but for collecting references. so i made one.
it's native, it's local, and it's blazing fast.
we analyzed >100k posts from r/ChatGPT over 3 years
on one hand, we saw ChatGPT quickly become normalized as an everyday consumer product, which is pretty cool
on the other hand…
Conservatives no longer trust the word "democracy." But they still believe in a functioning, citizen-led government. Fantastic new research unpacks what it is that they believe in, and why they think Trump is restoring the nation.
https://t.co/G59u1jMmDu
Humanity's ability to know, reason, judge, and act well is the foundation of science, democracy, crisis response, & management of AI itself.
AI poses serious risks to that foundation.
New paper on epistemic risks by 30 experts calls for attention to this. Link in thread.
I want some kind of LLM workflow tool.
• Ability to manage a set of input files (Markdown or similar), plus other general-purpose context.
• With real-time collaboration. (And maybe some concept of snapshots or VCS integration.)
• And the ability to create/manage a inference workflows and a stored set of prompts.
• Access to general-purpose coding agents (and not just chat models).
• Some concept of compiled outputs/inference results (which ideally can be shared externally).
Many projects have this feeling: "there is all this stuff, which I want to process/compute over in this iterated way, with some build artifacts being important/worth saving." GNU Autotools x Notion or something. Is anyone building this?
What could it mean for an AI to be "politically neutral”? And can we measure it? New paper + dataset.
We propose a defn that applies to any type of conflict: a neutral response should maximize approval on both sides of an issue, while keeping that approval balanced.
1/🧵
Excited to share our new work on Reinforcing Human Behavior Simulation via Verbal Feedback.
Can human simulators learn from feedback, not just rewards?
Most RL for LLMs turns feedback into a single score. But human behavior is rarely just right or wrong. It is social, contextual, subjective, and multi-dimensional.
A score can tell the model what is better. Verbal feedback can tell it why.
Meet DITTO + SOUL.
Paper: https://t.co/G0cEHr53h0
Code: https://t.co/6osJizwUDi
Model: https://t.co/yIAvpbKPSd
A new experiment involving 1,500 participants in 30 decision environments finds that AI advice depolarizes choices ~on average~, moving participants away from their initial leanings.
However, sycophantic AI increases polarization (p < .001). This poses a potential societal problem given that AI becomes increasingly sycophantic the more people engage with it and it customizes answers to match user preferences.
This suggests that the design features of AI are going to be critical to the impact is has on individuals, groups, and society. The technology can amplify or mitigate intergroup conflict, depending on how it's designed.
https://t.co/rcghZd1LGO
Math is easy* because it has verifiable outputs and few messy judgement choices to make.
Which AI labs have the guts to make advancing social science a priority? It may actually do more for human flourishing to unlock sociology, econ & psych reseach.
* For AIs, not for humans
User simulators have emerged as promising tools for building interactive AI, but what makes a “good” simulator?
We reframe the problem as what creates downstream value for humans
Our new simulator test: how an LLM assistant trained with the simulator performs with human users🧵
Spotted in the NYC subway. “Zero screen time.” An iPod Shuffle ad in 2026.
When we built the iPod, the goal was the technology disappeared and you could have your music wherever you were. 1,000 songs in your pocket.
Now we’re living through a moment where people are actively looking for ways to disconnect from the infinite feed, algos, and constant notifications. That doesn’t mean technology is bad. It means the best technology understands when to step back.
Not every problem needs another screen, another menu, or another layer of complexity. Constraints create freedom (read: @DavidEpstein new book Inside the Box). And often removing features creates a better product than adding them.
The future of technology shouldn’t just be more engagement. It should help us be more human.
"The whole stack should be completely adaptable, and should basically optimize on the fly to whatever task you have."
Co-founder @sarahookr in @TechCrunch on why we built AutoScientist.