Back to News
RSS feedarxiv.org

Benchmarking Automatic Prompt Optimization for Large Language Models with Chess

Summary

The paper introduces a chess benchmark for studying automatic prompt optimization (APO) on frozen large language models, whose weights remain unchanged while prompts are optimized. It is designed to address benchmark saturation, possible contamination of public test sets, and the cost of repeatedly evaluating prompts during optimization. The benchmark contains 1,118 Lichess puzzles and uses exact-match scoring, engine-based evaluation of alternative moves, adjustable difficulty, and a renewable supply of fresh problems. It also connects puzzle-solving performance with short game-play rollouts in the same domain, rather than measuring isolated puzzle answers only. The authors evaluate six APO algorithms on eight target models, examining baseline strength, responsiveness to optimization, and whether optimized prompts transfer across models and to game play. The strongest evaluated model, Gemini 3.5 Flash used as the meta-model, solves about 55% of the puzzles, leaving measurable room for improvement. Results distinguish gains, unchanged performance, and regressions across methods and models. The authors argue that chess is challenging, discriminative, renewable, and affordable for APO evaluation, with the full study costing around $800. They release the puzzles, optimization and evaluation code, and dataset-renewal scripts.