Conflicting Supervision Changes Model Commitment, Not Capability
Summary
This paper asks what happens when a model is trained on the same problems expressed through two incompatible but individually correct conventions. It argues that the learning-rate schedule is part of the mechanism determining which convention the parameters ultimately favor, rather than merely a background training setting. A theoretical bound separates the effects of data arrangement and schedule: arrangement enters through the period of repeated blocks, while the schedule determines how much weight can be placed on any single point in training. The authors compare ten orderings of one corpus under otherwise matched budgets and runs, repeating the experiments across learning-rate schedule families. With a constant learning rate, the allocation measure spans 0.2221 and the contrast floor reaches 11.63, increasing monotonically as the data become more blocked. Under the cosine schedule used by all published arms, the ten orderings resolve into two distinguishable states rather than ten, indicating that decay moderates the ordering effect. Across twelve decayed-schedule arms, the sum of the two convention-specific accuracies stays constant within 9.7%, while the allocation share ranges from 0.04 to 0.87. The authors describe this as a 12.29-sigma arrangement switch that is exactly zero under a convention-agnostic metric: the model’s overall capability does not change, but its commitment to one convention does. Marking the convention in the prompt removes the switch and reaches 87.5% of the union ceiling. The results therefore distinguish capability from path-dependent commitment and show why exact-match evaluation that ignores conventions can miss the effect.