Systematic generalization is the ability to solve new problems by recombining familiar atomic elements, but current evaluations often simplify this capability. They may use nearly linear action composition, productivity-based tests, or goals that explicitly identify the required actions. The authors introduce TranSGrid, a unified testbed that combines deductive, inductive, and abductive reasoning. Across 4,800 instances and seven Transformers, every model performed substantially worse on TranSGrid than on a held-out test set: the largest model reached 79.6% on the test set, compared with 55.3% on TranSGrid and 15.8% on its hardest subset. This gap remained within the models’ training-length range, indicating that productivity alone does not capture systematic generalization. Two modified TranSGrid variants separately restored either near-linear action composition or action-explicit goals, reducing the inductive or abductive demand; solve rates then returned to roughly held-out-test performance. The results suggest that existing tasks remove one or both of these reasoning requirements, and that a comprehensive evaluation must preserve deductive, inductive, and abductive demands together.
AI News
The latest AI releases, research, products, and industry updates.
Loading...