What is it about?
When companies test or train chatbots, they can't recruit thousands of real people for every experiment. So they ask one AI to pretend to be the human user and talk to another AI. But do those fake users actually sound like real people? We built MirrorBench, an open-source tool that answers this. It compares AI-generated user messages against real human conversations from four public datasets, scoring them on word variety and on whether other AI models — and human annotators — can tell the difference.
Featured Image
Photo by Jason Leung on Unsplash
Why is it important?
Simulated users now quietly underpin how AI assistants get tested and trained, yet almost nobody checks whether the simulation is any good. If the fake users are unrealistic, every result built on them inherits that flaw. MirrorBench is the first benchmark to isolate the user's side of a conversation and judge it purely on human-likeness, separate from whether the task succeeded. It also exposes a trap: swap the AI acting as grader and the rankings can change — so single-judge results shouldn't be trusted on their own.
Perspectives
The finding that stuck with me was how differently things look depending on which lens you use. Models that judges rated as the most convincingly human often used noticeably less varied language than real people did — realism and diversity pulled apart rather than moving together. That tension shows up again and again across datasets, and it's a reminder that "sounds human" and "behaves like a human population" are not the same measurement. The other result worth flagging for practitioners: a small open-source model fine-tuned specifically for user simulation held its own against much larger frontier models. Specialization may be a cheaper path to realistic user simulation than scale. MirrorBench is open source at github.com/SAP/mirrorbench, and I'd genuinely like to see people break it — especially in non-English and clarification-heavy settings, where our results show proxies struggle most.
Ashutosh Hathidara
SAP Labs LLC
Read the Original
This page is a summary of: MirrorBench:
A Benchmark to Evaluate Conversational User-Proxy Agents for Human-Likeness, August 2026, ACM (Association for Computing Machinery),
DOI: 10.1145/3770855.3817496.
You can read the full text:
Contributors
The following have contributed to this page







