Back to News
RSS feedarxiv.org

Evaluating AI Agent Resilience Under Accumulating Challenges

Summary

This paper argues that sustained deployment of generative AI agents requires more than succeeding at isolated tasks. It proposes two complementary evaluation concepts: operational resilience, covering recovery from blocked work, preservation of progress, and communication of limits; and considerate participation, covering adaptation to affected people, role boundaries, and the surrounding workflow. The authors study 120 simulated healthcare trajectories involving two generative AI models and 12 stakeholder-derived tasks under light, medium, and heavy challenge. They compare textual action plans, prompted internal assessments, and quantitative reports of workload and affect as difficulties accumulate. In the resilience analysis, agents move from self-directed recovery toward greater reliance on humans. Structured reports show rising workload and negative affect, but textual responses rarely explicitly express strain. In the participation analysis, agents expand beyond task-focused adaptation toward reframing tasks, attending to others, adjusting role boundaries, and coordinating more broadly, with different patterns between actions and internal assessments. The paper distills five deployment dilemmas concerning persistence, attention, role boundaries, state disclosure, and escalation. It argues that these dilemmas require stakeholder specification and can guide future work on learning, situated evaluation, and embodied adaptation.