Back to News
RSS feedarxiv.org

More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses

Summary

Automated generation of LLM harnesses is intended to improve inference by assigning different programs to different tasks, but repeated execution of the same program can also increase the chance of finding an answer. This study introduces a controlled evaluation that separates raw answer coverage, repeatable task-specific advantages, and gains from selecting a harness before execution. On 386 MATH-500 tasks, the authors compare eight generated harnesses and a baseline with nine byte-identical copies of the baseline, running each member three times. Even identical programs produce 2.16 percentage points of repeat-averaged oracle headroom, showing that coverage can increase without program specialization. Generated programs have more repeatable score patterns, but those patterns mainly expose persistent weaknesses: the generated programs lose to the baseline on all three repeats for 100 tasks, whereas persistent wins occur on only one task and depend on answer extraction. A frozen selector adds 0.00 percentage points. After 27 harness executions, both the generated and identical-copy populations reach 98.70% oracle coverage. The results therefore do not establish stable complementarity at three repeats. BIRD traces attribute failures to mechanism implementation, activation, and output validity. The authors conclude that coverage and repeatability alone are insufficient evidence of useful specialization. They propose that harness diversity should be judged by whether task advantages persist across executions, support usable decisions, and outperform additional executions of a fixed program under matched inference budgets.