VAmoS Part Deux Introduces a More Realistic Voice-Agent Benchmark
Summary
The paper introduces VAmoS Energy, a benchmark designed to evaluate production voice agents under conditions that combine multiple requests, background speech, and impatient callers. It contains 100 simulated utility-billing and payment-assistance calls, with each caller making two to four requests. Agents can use 16 tools connected to a stateful Stripe billing twin and the Apache Fineract loan engine, but account access is blocked until caller verification succeeds. The benchmark uses public household electricity data and a policy based on Pennsylvania residential billing rules. An LLM-based verifier checks both the agent’s actions and the figures it speaks against explicit requirements; in calibration, it agreed with a code verifier on 99.1% of checks. Across 14 voice stacks, with each task repeated three times, completion rates ranged from 17.3% to 44.7%. Grok Voice ranked first, while Gemini 3.8 Live and GPT-Live followed at roughly the same cost per call. Adding background television reduced pooled completion from 38.7% to 8.6%. The simulated caller sometimes accepted incorrect outcomes because it could hear the agent’s words but could not inspect its actions. The authors therefore argue that voice-agent evaluation must cover the entire call, including speech, state-changing actions, verification, and performance under competing audio.