Back to News
RSS feedarxiv.org

Defining AI Agents Through Criteria, Metrics, and Benchmarks

Summary

The survey addresses the lack of a standard definition for “agent” in artificial intelligence, which makes agent research harder to evaluate, compare, and reproduce. It organizes agenticness around five dimensions: interaction with an environment, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, the authors review how prior research has described the capability and synthesize the metrics, benchmarks, and evaluation frameworks used to measure it. The review identifies established evaluation approaches as well as areas where methods remain limited or inconsistent. The authors also introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods collected in the survey. Together, the survey and compendium provide a shared structure for comparing capabilities across artificial agents, with the stated aim of supporting clearer communication, more reproducible research, and more systematic study.