SAGE Uses Statistical Acceptance Gates to Make Self-Evolving Agents Safer
Summary
Large language model agents can self-evolve by editing persistent skill documents that store workflow, tool-use, and decision rules. Existing gates typically accept any edit that improves an aggregate validation score, which can permanently regress items the agent already solves and can select an apparently best edit because finite noisy validation scores are upward biased. SAGE addresses these problems with per-item paired comparisons: the current and edited skills are tested on the same validation items, exposing regressions and penalizing them asymmetrically. It also applies a one-sided paired statistical test and commits an edit only when wins are reliable relative to losses; otherwise, it abstains. The method is conservative and exactly recovers the baseline at a boundary setting, so it filters only a subset of the baseline’s accepted edits. Under an equal-budget evaluation across five benchmarks and four backbone LLMs, SAGE reduced the regression rate in 19 of 20 settings and matched the baseline in the remaining setting. With DeepSeek-V4, regression rates fell from 36.5% to 0% on LiveMath and from 42.8% to 0% on OfficeQA. SAGE also achieved the highest final score in all 20 settings, including an increase on LiveMath from 34.15 to 48.78.