ShowTellArena Benchmarks AI Agents’ Understanding of Narrated Business Workflows
Summary
ShowTellArena introduces a benchmark protocol and public dataset for evaluating what an AI agent understands after watching a narrated business demonstration. Version 1.0 contains 50 business workflow tasks, recordings, screenshots, narration, fixture seeds, and 502 questions covering finance, hiring, procurement, customer decisions, inventory, and logistics. The protocol keeps the business scenario and quiz fixed while allowing each product to capture the lesson through its own teaching interface. Its questions probe operational rules, boundaries, exceptions, and errors in proposed automations. The authors analyze 218 selected pilot attempts across 39 workflow cases, including 28 cases attempted by all three evaluated systems. The pilot revealed both incorrect answers and failures to complete the teaching experience. The release also documents verification gaps, uneven coverage, exclusions, and the provenance of grading. The authors present the dataset and assessment workflow as inspectable resources that others can extend, and explicitly state that the selected pilot is not a controlled ranking of products.