Current chatbot interfaces present AI models as authoritative and logical oracles. This makes it difficult for users to calibrate their perception of the model's capabilities, and often incentivizes users to defer to AI judgment.
In our position paper at the Pluralistic Alignment Workshop @ ICML 2026 @pluralistic_ai, we give this pattern a name: deference by design.
https://t.co/snfiVltl6I
🧵
If AI is going to participate in our thinking, we need to build systems that keep humans in charge of the important choices.
I built Priori, an interface for human-AI collaboration, toward that goal last year.
I prototyped the early versions of Priori last year and presented at Cosmos & FIRE's AI x Truth-seeking Symposium in December. Since then, multi-agent orchestrators and agentic AI have come far, and some lower-level choices need less auditing than they used to. If anything, that has made the underlying question more urgent: as systems grow more capable, human judgment needs to become easier and more worthwhile to express.
Thanks to @cosmos_inst and @TheFIREorg for supporting this work, and to @sebkrier@glukianoff@jon_rauch@AndrewMayne@PhilippKoralus for the thoughtful feedback at the first symposium. I will have more work on this direction coming soon :)
This looks socially very valuable, as a way to retain meaningful human agency while working with AI, but on top of that just a really good UI feature.
There are plenty of coding projects I’ve been doing where I’ve only figured out several hours in that the model has baked in a wrong assumption about how to fix the problem from the outset. Eliciting that early doors would have saved loads of time.
We also may not need mechanistic interpretability to make progress on (user-facing) transparency. Recent work like OpenAI’s "confessions" suggests that some of the infrastructure for surfacing decision-relevant information may already exist inside major labs.
One next step may simply be to expose more of that structure to the user. Personally, seeing the key decisions a model made would on average be much more useful than checking 20,000 COT tokens, especially in CLIs. For example:
If AI is going to participate in our thinking, we need to build systems that keep humans in charge of the important choices.
I built Priori, an interface for human-AI collaboration, toward that goal last year.
@cosmos_inst@Brendan_McCord@lawhsw@Houda_nait Waiting until everything feels finished can mean waiting too long. So I'm going to try sharing more of what I'm working on. Follow if you're interested in human-AI collaboration, scalable oversight, and human autonomy :)