AI evaluation is entering an interactive benchmark era.
Across tool-use agents, web/OS benchmarks, multi-agent systems, and reliability evaluations, interaction is becoming central to how modern AI systems are tested.
But the field risks adding interaction faster than it develops the scientific principles for evaluating interaction.
Our position:
Interactive evaluation is not just longer tasks, tool use, or multi-turn interaction.
It requires a design science for mapping trajectories to valid evaluative claims.
๐ https://t.co/lKGuDOuBZy
๐ป https://t.co/LkadiPYnnw
๐ฉ๐ผโ๐ป Real or Robotic? ๐ค Can LLMs accurately simulate qualities of human responses in dialogue?
Human conversations with LLMs are great for assessing the capabilities of LLMs. But having lots of folks chat with LLMs is challenging (๐ฐโณ๐ต๏ธ). Could we have another LLM *simulate* being a human talking to an LLM as a substitute?
In our new preprint, we test whether models can roleplay as the human in human-LLM conversations. Using the WildChat dataset and 100K+ simulations we test how well these LLM responses actually mimic with human ones. Our study spans ๐ฌ๐ง English, ๐จ๐ณ Chinese, and ๐ท๐บ Russian, using 21 linguistic metrics like lexical, semantic, syntactic, and stylistic features.