Back to News
RSS feedarxiv.org

Executable Checks Improve Some Scientific Coding Agent Repairs

Summary

The Rules to Tools (R2T) approach gives scientific coding agents callable checks derived from public requirements such as equations, boundary conditions, and expected outputs. In matched SciCode repair groups, the tool condition completed 29 of 30 repairs, compared with 26 of 30 for written checks alone. In an eight-task cohort, checks led 15/16 to 13/16, but the task-cluster bootstrap 95% interval for the difference was [-12.5, 43.75] percentage points, so the aggregate advantage remains uncertain. A larger shared-definition cohort tied at 13/24 per group, and five development-exposed tasks instead scored 7/10 with text versus 3/10 with tools. The task-level pattern was mixed: checks favored tasks 17, 77, and 11, while text favored task 37. A fresh source-through-Python condition also reached 15/16, matching the dedicated command aggregate. In a matched PDE comparison, checks scored 24/24 versus 23/24 for detailed text and produced 31.2% less reported model output. Output savings varied by cohort, while public CPU use increased in both task-ID cohorts. The results therefore support task-dependent effects and agent-side cost changes, rather than a universal benefit from executable checks.