Back to News
RSS feedarxiv.org

EmailBench Benchmarks LLM Agents on Enterprise Email Tasks

Summary

Enterprise email agents must retrieve information, change structured state, reason over time, and coordinate multiple steps. EmailBench introduces a self-contained benchmark with 206 email and productivity scenarios spanning 16 task categories. It combines a typed email API with provider-neutral names, a deterministic synthetic corpus inspired by Enron email, and scenarios informed by aggregate task-intent telemetry from an interactive prototype. Evaluation uses 258 executable static assertions together with 211 LLM-based rubrics. The authors test eight language-model configurations on a fixed single-user corpus. The strongest configuration passes only 33.5% of scenarios, even though 99.7% of its tool calls complete without an observed API failure. Pass rates differ substantially across task categories, showing that valid tool execution does not necessarily produce correct end-to-end task completion. The benchmark is intended to support controlled evaluation of email agents; broader tool coverage, multi-persona testing, and repeated-run evaluation are identified as future work.