DuplexSpeechBench-IFEval Evaluates Implicit Instruction Following in Full-Duplex Voice Agents
Summary
The paper introduces DuplexSpeechBench-IFEval (DSB-IFEval), a benchmark for testing whether full-duplex voice agents can infer conversational behavior from roles or personas rather than relying only on explicit turn-management instructions. The benchmark contains 1,038 test cases covering eight assistant roles and five conditioning protocols: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona-and-rule conditioning, and instruction conflict. It evaluates real-time floor management with a deterministic Instruction Adherence Score (IAS) and persona-consistent content with an LLM-judged Persona Adherence Score (PAS). Across six real-time speech systems, the results show architecture-dependent trade-offs. F-Actor and PersonaPlex are more sensitive to implicit conditioning, with adherence falling 9.7% and 4.5%, respectively, under persona-only instructions. GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat follow persona-consistent content strongly, but their floor behavior changes little between explicit and persona-only instructions and remains limited on several proactive actions. The systems also struggle to override persona directives when they conflict with safety requirements. The authors conclude that inferring role-implied behavior, executing it at the right conversational moment, and resolving competing instructions are separate challenges for full-duplex voice agents.