Back to News
RSS feedarxiv.org

Preregistered Study Finds Evidence Masking Improves Compositional Generalization

Summary

A preregistered confirmation study tests whether restricting what modules can read improves what a system learns to compute. The researchers evaluate 60 four-cell systems that share a frozen language-model backbone and communicate through learned continuous packets. They compare five conditions involving evidence masking, ownership markers, and neutral replacement of foreign evidence, across six initialization clusters, two data orders, and one new task world. When markers were available, masking improved held-out two-operation and three-operation compositional accuracy by median paired differences of 0.846 and 0.859, respectively; all 12 paired comparisons met the required margins, and the full preregistered behavioral criterion passed. An unmarked replication also passed. However, no globally visible system passed the marker-following check, leaving the role of usable role information unresolved. The neutral-filler condition produced seven full generalizers, but its decomposition tests were inconclusive. Packet interventions in all 18 audited masked systems changed predicted intermediate values on eligible cases, although the finite, success-conditioned audits do not establish mediation. The authors conclude that the tested masking regime has a large benefit, while its detailed causal explanation and broader generality remain open. Protocols, results, and checkpoints are public.