GenUI-Harness Trains Agents to Generate Data-Aware Interfaces
Summary
Text-based interaction can create cognitive overload, ambiguity, information clutter, and slow input for complex tasks. The paper introduces GenUI-Harness, a multi-agent system that combines a Tool Agent for information retrieval and task execution with a GUI Coder Agent that resolves ambiguities and generates front-end code for structured interfaces. It addresses two reinforcement-learning problems: executing generated UIs to obtain verifiable rewards is expensive, while LLM-as-a-Judge rewards can be exploited through reward hacking. Dynamic UX is a lightweight sandbox package for dynamic interaction and reward collection, and Reward Auditor is a meta-reward mechanism that monitors reward distributions and distills diagnostic patterns into a shared rubric and scoring specification. The authors also introduce UI-TAU Bench, a benchmark for active human-agent interaction through generated UI code. It uses 10 real-world domain databases built from public data and follows Tau-Bench tool-use settings, with Lite and Full splits containing 300 and 1,000 tasks. On Lite, GenUI-Harness improves average Pass@3 by 4.48 percentage points over smolagents. Training raises a 4B backbone’s Pass@3 from 9.33% to 58.00%, exceeding the reported 46.67% for Claude Opus 5 in this evaluation. The harness remains robust on ambiguous and non-ambiguous queries. In a reviewer survey, generated interfaces reduce average dialogue rounds from 3.4 to 1.2, although the evidence is limited to the evaluated database-backed workflows.