Back to News
RSS feedarxiv.org

AREX-2 Advances Self-Improving Agents with Long-Horizon Reflection

Summary

AREX-2 studies how LLM agents can improve their own solutions at test time through repeated iteration. The authors separate this ability into reflection, which produces a better solution than the current one, and long-horizon execution, which keeps the process effective across many rounds. They hypothesize that both capabilities are domain-agnostic, so they synthesize improvement trajectories from machine learning and algorithmic programming tasks where feedback can be verified. An agent built on Qwen3.8-27B and trained on this data scores 81.8 on MLE-bench Lite and 70.7 on Frontier-CS. It also transfers to deep-research evaluations, scoring 84.0 on BrowseComp, 52.6 on HLE, 92.2 on GAIA, and 93.8 on DeepSearchQA. The agent continues to improve when given a larger round budget, which the authors present as evidence that long-horizon reflective data can support self-improving agents beyond the domains used for supervision.