StoreBench Introduces a Live-Commerce Testbed for Autonomous Operator Agents
Summary
The paper introduces StoreBench, a live-commerce environment for evaluating and training autonomous operator agents that manage a mid-size online apparel store. Unlike static agent benchmarks, its customers order continuously, suppliers can reprice or fail, and market shocks may occur with little or no warning. Agents use the same 29 merchant tools as human operators and work under a windowed action budget, making simulated time depend on actions rather than model latency. Pass thresholds are calibrated against scripted anchor policies, rewards are hardened against documented reward hacks, and episodes replay identically for the same action sequence. The authors evaluate seven frontier LLMs across 11 scenarios lasting 30 to 45 days, plus a full simulated year, using three world seeds and matched reasoning effort. No model reaches the scripted smart-triage policy’s average performance: the strongest result, from DeepSeek-V4-Pro, passes 49% of task-seed cells versus 97% for the heuristic. Human experts also outperform every model on the mean composite score, 0.708 versus 0.700. Longer simulated runs improve most models, while GRPO training raises Qwen3.5-27B’s held-out mean composite from 0.136 to 0.373 after training on five disjoint tasks. The release includes example training tasks, sample trajectories, and scoring and verification tools, but withholds the full environment and evaluation suite to limit benchmark contamination.