What happens when you compare the distributions of real and simulated user behaviors?
🔍 The gap is large.
We introduce a method to measure this gap and evaluate 24 LLM-based user simulators across coding and writing tasks.
@convai_uiuc@MSFTResearch@berkeley_ai
🧵 1/N