Mathematical Transfer in LLMs Depends More on Reasoning Approach Than Topic
Summary
This study examines whether mathematical training data should be selected by topic or by reasoning approach. It compares same-approach (SA) sources, which use the target problem's method in a different topic, with same-topic (ST) sources, which share the topic but use a different method. Two counterbalanced 2x2 designs cover probability and combinatorics with invariant and double-counting reasoning across 2,000 problems, and number theory and geometry with complement and pigeonhole reasoning across 800 problems. Each cell serves as a held-out target in turn, and the design balances source roles so additive source-quality effects cancel in the aggregate comparison. Across five base models and three training seeds per design, SA outperformed ST in all 40 seed-pooled model-target comparisons. In the primary design, model-level gains ranged from 8.2 to 16.2 percentage points, averaging 10.8; in the second, they ranged from 12.0 to 16.0 points, averaging 14.3. All ten model-level 95% confidence intervals excluded zero. ST examples were nevertheless more similar to targets under embedding and lexical measures, so the observed advantage was opposite to statement-level resemblance. The results support reasoning approach as a more effective matching criterion than topic for the evaluated mathematical combinations.