Back to News
RSS feedarxiv.org

Self-Evolving Agent Harnesses Across Multiple Tasks

Summary

A harness is the code surrounding a language-model agent that manages prompts, tool calls, context, and execution. This study proposes a recursive self-improvement framework in which the same frozen model, running the same harness version, first solves tasks and then acts as a proposer that reads complete run records and edits the harness directly. Each evolution batch uses tasks from five benchmarks spanning different domains, with training and held-out tasks separated and five additional out-of-distribution benchmarks reserved for evaluation. The authors describe the process as two-stage training: multi-task pretraining followed by continual training. Starting from a 49-line seed harness, the first stage raises average performance by 4.48 points on in-distribution benchmarks and 12.64 points on out-of-distribution benchmarks, surpassing Codex on the former and matching it on the latter. Continued evolution on the out-of-distribution Claw-Eval benchmark increases its score from 66.17 to 68.06, exceeding Codex. Analysis identifies output truncation, history compaction, and independent review as mechanisms that emerged during evolution.