RSS feedarxiv.org
GMA Benchmark Tests AI Agents on Complex Mobile Workflows
Summary
Researchers introduce GMA, a benchmark for evaluating general mobile assistants in realistic, challenging scenarios. It spans seven open-source applications and 300 tasks across four difficulty tiers. Tests of eight frontier models show sharply declining performance as workflows become more complex. Ablation studies find that context retention and explicit state tracking can improve results, although effective harness designs vary by foundation model.