Tau-Elicitation Benchmarks Multi-Turn Entity Extraction in Voice Agents
Summary
The paper introduces Tau-Elicitation, a 200-task benchmark for testing whether voice agents can collect names, addresses, identifiers, dates, and times exactly across 10 entity types, controlled difficulty levels, caller realisms, and three environments. A matched text agent completes every task, while four voice configurations reach robust exact success rates of only 0.14 to 0.41. Voice agents verify more often when entities are difficult or unfamiliar, and sometimes after incorrect captures, but they do not increase verification for the caller voice on which they perform worst. Only 24% to 37% of verified errors are successfully repaired. A scaffold requiring spelling, read-back, correction, and confirmation improves robust Pass^3 by 14 to 31 percentage points, but adds 21 to 28 seconds per call. Spelling variations and restarts do not measurably change exact success, whereas mispronunciation increases repair effort. The results identify strategy selection and reliable recovery as the main bottlenecks in exact spoken entity collection.