Back to News
RSS feedarxiv.org

WinSyn Automates Realistic Evaluation for Enterprise QA Agents

Summary

Enterprise question-answering agents must work across evolving and sometimes conflicting emails, chats, documents, and other workplace artifacts. The WinSyn paper introduces an automated pipeline that generates synthetic email datasets representing realistic workplace scenarios, together with short- and long-form questions and data-grounded gold answers. The simulated projects run for several months and can involve up to 25 employees in different roles, emphasizing ambiguity, distributed evidence, and naturally occurring queries. The authors evaluate several standard agentic baselines with current frontier models. Average aggregate scores remain below 80% on every dataset, indicating substantial room for improvement. The results support using higher-complexity, more realistic evaluation data when developing enterprise deep-research systems and assessing their readiness for deployment.