EmailBench Benchmarks LLM Agents on Enterprise Email Tasks
Summary
Enterprise email agents must retrieve information, change structured state, reason over time, and coordinate multiple steps. EmailBench introduces a self-contained benchmark with 206 email and productivity scenarios spanning 16 task categories. It combines a typed email API with provider-neutral names, a deterministic synthetic corpus inspired by Enron email, and scenarios informed by aggregate task-intent telemetry from an interactive prototype. Evaluation uses 258 executable static assertions together with 211 LLM-based rubrics. The authors test eight language-model configurations on a fixed single-user corpus. The strongest configuration passes only 33.5% of scenarios, even though 99.7% of its tool calls complete without an observed API failure. Pass rates differ substantially across task categories, showing that valid tool execution does not necessarily produce correct end-to-end task completion. The benchmark is intended to support controlled evaluation of email agents; broader tool coverage, multi-persona testing, and repeated-run evaluation are identified as future work.