Excited to share what we have been building at @MIT
A new way to evaluate social simulation: Social Simulation vs. The Real Future. Each week, simulators predict how people will think, act, and change. Predictions are locked in before the results come out. We then score them and publish the scores.
🏆 Building a simulator? Bring it: https://t.co/iuSvSlbtZh
🔑 Want to build the arena with us? It is open source: https://t.co/pjwtIeRQpQ
Social simulation is now a product. Evaluation still has no common standard.
What if we tested every simulator against the real future?
We built Social Simulation Arena with researchers from MIT, Stanford, Harvard, CMU, UC Berkeley, and beyond.
Persona-prompted populations. Digital twins. Silicon samples. Synthetic populations. Worlds of a billion agents.
We have more ways than ever to simulate human behavior. But we still lack a standard everyone can trust to tell which ones work.
Most studies use their own datasets, often historical surveys whose answers are already online. A model that has seen the answer sheet can look brilliant.
Social Simulation Arena is a different kind of test: prospective, independent, and shared.
Simulators submit and lock their forecasts before each data release. When the real results arrive, every entrant is scored by the same rules.
Can a simulator predict a public before it moves?
If you’re building one, bring it.
https://t.co/SIytJoXeIR
#MIT #Stanford #Harvard #AI #Agents #LLM #SocialSimulation #Evaluation
So what does “simulating thought” actually look like?
Try do(Transparency) yourself: https://t.co/ydAxb4QVF2
In our demo, a person is represented through an inspectable belief graph. Change one belief and watch the consequences propagate through the rest of their reasoning.
The point is not that a mind is literally a graph. The point is that the reasoning behind a simulated behavior should be something we can inspect, intervene on, and test.
tl;dr: Most LLM social simulations are still black boxes: demographics in, behavior out.
We argue the field needs the same conceptual shift psychology once made in the cognitive revolution: from behaviorism to cognitivism. That means moving beyond plausible outputs toward agents with (1) inspectable reasoning, (2) grounding in real individual experience, and (3) consistent counterfactual reasoning.
Our paper, Simulating Society Requires Simulating Thought, lays out where current simulations fall short, how cognitively grounded agents can be built, and how to evaluate reasoning fidelity rather than output plausibility.
Paper: https://t.co/siyNyvmlmm
Website: https://t.co/ydAxb4QVF2
Joint work with Jiayi Wu, Zhenze Mo, @ao_qu18465, Yuhan Tang, @kyzhao_ivy, @yule_gan, Jie Fan, Jiangbo Yu, Jinhua Zhao, @pliang279, Luis Alonso, and @Larson_MIT. A @medialab project with collaborators at MIT, Northeastern, Brown, and McGill.
Presented at NeurIPS 2025 (Position Paper Track).
Behavioral agreement is not sufficient evidence of simulation fidelity.
When the question is about belief formation, individual differences, or interventions, the same observed behavior can arise from very different underlying mechanisms. An agent can give the right answer for the wrong reason.
We should therefore validate not only what an agent says, but how its beliefs are formed and how they change.
That requires the same move psychology once made: from behaviorism to cognitivism.
Want to see how hard this actually is?
Try the quiz: https://t.co/LBOARf6tNx
We put four real items on the project page. You read a stranger's interview excerpt, predict their answer, and then reveal what they actually said.
The models in the paper average around 75% on beliefs, but only 58 to 69% on updates.
And a huge thank you to the 54 people who sat down and thought out loud for us.
tl;dr: Give an LLM a person's own words, and it can recover what they believe surprisingly well. But predicting how that same person will update their beliefs is much harder. More transcript or memory doesn't close the gap.
Our paper, HugAgent: A Human Simulation Benchmark for Individual-Level Reasoning, takes on this harder problem: simulating how a specific person reasons and updates, rather than simply predicting their next answer. We collect think-aloud interviews with an LLM-driven chatbot and test models on beliefs, reasoning, and counterfactual belief updates.
Paper: https://t.co/gWZq8wAyeK
Website + 4-item quiz: https://t.co/LBOARf6tNx
Code and data: https://t.co/SQP0BJYjvz
Interview chatbot: https://t.co/8ofHoxhhA0
Joint work with Zhenze Mo, Yuhan Tang, @ao_qu18465, Jiayi Wu, @kyzhao_ivy, @yule_gan, Jie Fan, Jiangbo Yu, @hjian42, @pliang279, Jinhua Zhao, Luis Alonso, and @Larson_MIT. A @medialab City Science project with collaborators at MIT, Northeastern, Brown, and McGill.
Accepted at @emnlpmeeting 2026 (main conference).
Humans are equally consistent on both tasks two weeks later: 84.8% on belief states and 85.7% on belief updates.
Models aren't.
They reach 76 to 78% on belief states, but only 58 to 69% on updates:
Claude Sonnet 4.5: 68.6%
GPT-5.1 (reasoning): 67.1%
GPT-5-mini: 58.2%, roughly the majority baseline.
Each prediction is grounded in a real person's own interview and held-out responses.
Want to see how hard this actually is?
We put four unedited items on the project page. You read a stranger’s interview excerpt, predict what they’ll say next, and then reveal their actual answer.
The models in the paper average around 75% on beliefs, but only 58 to 69% on updates.
Try the quiz: https://t.co/LBOARf6tNx
And a huge thank you to the 54 people who sat down and thought out loud for us.
AI Agent Personas should simulate the structure of human reasoning.
I’ve been arguing that you cannot "invent" a digital expert agent using just prompt engineering. You have to extract the expert via deep interviewing.
A new NeurIPS paper, "Simulating Society Requires Simulating Thought" reinforces everything we've discussed about why thin, synthetic LLM personas fail.
Most AI agents operate as "behaviorists." When you prompt an LLM to "act like a senior economist," it relies on surface-level correlations from training data. It generates text that sounds expert-like, but lacks any internal belief structure.
1. Logical Inconsistency:
Without an internal model of how beliefs are formed, agents support a policy in one context but oppose it in another. The paper calls this "intervention-invariance mismatch" - beliefs don't update coherently when assumptions change.
2. Illusion of Consensus:
In multi-agent simulations, LLMs converge toward the median view (even more positive emotions as the other paper mentions) of the training data. They agree not because of shared reasoning, but because their statistical priors push them toward the center. Your expert's contrarian, hard-won perspective gets averaged out.
3. Identity Flattening:
LLMs reproduce stereotypical portrayals that erase intersectional variation. "The rich, positional knowledge of real-world stakeholders is replaced with monolithic, decontextualized simulations."
To fix this, we have to move from simulating speech to simulating reasoning. The authors propose a "Cognitive Modeling" approach.
"beyond output-level alignment toward aligning the internal reasoning traces of generative agents."
Their solution is SEMI-STRUCTURED INTERVIEWS to extract what they call "cognitive motifs" - minimal causal reasoning units that capture how a specific person actually thinks.
This is exactly why we built an interviewer system instead of a persona generator. You have to extract their actual belief structure through conversation.
Instead of predicting the next word, the agent must possess "Reasoning Fidelity", a structured map of beliefs, causal logic, and cognitive motifs.
How do you get this map?
You can't prompt for it. You have to interview for it, with AI. The paper explicitly validates the architecture we’ve built: using semi-structured interviews to elicit "causal explanations" and "reasoning traces".
This confirms why our Interviewer + Note-Taker multi-agent system is critical.
- The Interviewer builds the "Peer Status" necessary to get the expert to open up.
- The Note-Taker (the cognitive layer) extracts the "Cognitive Motifs", the distinctive logic blocks that define how that specific expert solves problems.
We are moving beyond the era of "acting like an expert" to Generative Minds; agents that embody the positional individuality and causal logic of the people they represent.
If you're building AI agents for strategy, decision-making, or stakeholder modelling, start by interviewing the human aspects of your agents.
Everyone "knows" that LLM agents are getting smarter.
They speak better. They argue better. They sound more human.
But I’m about to ruin a "fun fact" you probably believe:
Most of that "intelligence" is a mirage. And MIT has the research to prove it. 🧵👇
We are currently stuck in a "Demographics In, Behavior Out" paradigm.
We tell an AI: "You are a 40-year-old voter."
It spits out: "I care about taxes."
It looks real. But it’s actually just behaviorism on steroids. It’s mimicry, not mind.
Here is the scary part: The "Traceability Gap."
When an AI explains why it made a decision, it’s usually lying.
It’s called "post-hoc rationalization." The model makes a choice based on math, then invents a plausible-sounding story to justify it afterwards.
This leads to the "Illusion of Consensus."
Because models are trained to be "average," multi-agent simulations tend to agree with each other way more than real humans do.
They flatten the messy, intersectional reality of society into a boring, statistical mean.
So, how do we fix it?
A new paper proposes a shift: Generative Minds (GenMinds).
Instead of just predicting the next word, we need agents that build Causal Belief Networks.
Think of it as giving the AI a "mental map" of cause and effect.
With GenMinds, an agent doesn't just say "I support policy X."
It builds a graph:
Transparency → Reduces Crime → Increases Safety.
If you change the input (Intervention), the belief updates logically.
It’s not just text generation. It’s simulated thought.
The researchers also introduced RECAP.
It’s a new benchmark that stops grading AI on "does this sound good?" and starts grading on "does this make sense?"
Traceability. Adaptability. Consistency.
If we want to use AI for policy, voting simulations, or social science, we have to stop settling for surface-level mimics.
We need agents that can reason, not just recite.
Read the full paper here:
If you want to see how powerful data + maps are to tell a story, check out this powerful video and maps created by #SpatialEquityNYC. These maps allows users to use Census data to demonstrate the consequences of inequity. #mondaymapday#builtwithmapbox https://t.co/shdpgDwAag.
Hi everyone! If you are interested in what we are doing and joining our community, please attend our first town hall meeting in Jan. 25th. Check the details in the poster: