Teaching Agents to Code Reliably
Summary
The paper argues that reliable autonomous coding depends on three trainable behaviors: exploring different repository locations, producing diverse edits, and verifying patches against a reverted version of the original tree. It identifies three failure modes in current inference-time search: repeated attempts revisit the same location, different methods that could solve complementary issues are underused, and tests written by an agent for its own patch can accept incorrect repairs. A search procedure guided by execution feedback and a verifier that scores each patch against its reverted tree resolves 52.8% of SWE-bench Verified issues while using 48.1% of the agent-steps required by an eight-sample baseline. The researchers then train these behaviors into the policy rather than relying on an external scaffold. On 270 issues held out from supervised fine-tuning and reinforcement-learning training, weighted supervised fine-tuning increases pass@1 from 31.9% to 35.2% and pass@8 from 46.7% to 51.1%. A reinforcement objective trains the verifier with gold-labeled repairs and incorrect variants, rewarding assertions that distinguish them; pass@1 rises to 43.0%, pass@8 to 60.7%, and verifier precision from 26.8% to 41.7%, while false acceptance is reduced by more than half. Resolution improves on two of three out-of-distribution suites, and verifier precision improves on all three. The gains remain present across 7B, 14B, and 30B models compared with published coder baselines.