Back to News
RSS feedarxiv.org

What Current Systematic Generalization Tasks Miss: A Reasoning-Centered Analysis

Summary

Systematic generalization is the ability to solve new problems by recombining familiar atomic elements, but current evaluations often simplify this capability. They may use nearly linear action composition, productivity-based tests, or goals that explicitly identify the required actions. The authors introduce TranSGrid, a unified testbed that combines deductive, inductive, and abductive reasoning. Across 4,800 instances and seven Transformers, every model performed substantially worse on TranSGrid than on a held-out test set: the largest model reached 79.6% on the test set, compared with 55.3% on TranSGrid and 15.8% on its hardest subset. This gap remained within the models’ training-length range, indicating that productivity alone does not capture systematic generalization. Two modified TranSGrid variants separately restored either near-linear action composition or action-explicit goals, reducing the inductive or abductive demand; solve rates then returned to roughly held-out-test performance. The results suggest that existing tasks remove one or both of these reasoning requirements, and that a comprehensive evaluation must preserve deductive, inductive, and abductive demands together.