Back to News
RSS feedarxiv.org

Diagnosing Why Terminal-Agent Training Stalls

Summary

Using a frontier model such as Claude Opus as a meta-agent to generate terminal tasks and verifiers is becoming common, but a runnable Docker image and executable test suite do not ensure a faithful training pipeline. This study identifies three failure classes: invalid benchmarks, brittle evaluation harnesses, and reward misalignment. Redesigning prompts and extending context increases baseline solvability by 5.6 times. However, a 9B model reaches a mean pass@2 of 81.3% within 20 steps on tasks generated by Claude Opus, indicating saturation under that setup. Adding harder tasks drops mean pass@2 to 20.6% without changing the training configuration, showing that the useful solvability range depends on the model. The authors argue that reliable meta-agent pipelines should treat solvability-band calibration, verifier audits, and accounting for infrastructure errors as primary evaluation criteria rather than post-hoc diagnostics.