Witness Overlap Infers Directional Provenance in Open-Weight Model Families
Summary
Open-weight models are frequently fine-tuned, aligned, merged, and re-released, but many provenance audits can establish relatedness without determining which checkpoint came first. The paper introduces Witness Overlap, a prompt-free, training-free white-box test that adds a third checkpoint from the same model family as a witness. It compares the local geometry around the two candidate endpoints and identifies the checkpoint that behaves more like a branching parent. On 176 LLM checkpoints spanning 16 families, the one-witness test oriented 95.3% of parent-child decisions when using Frobenius cosine. The authors also evaluated root identification, sibling discrimination, chain-structured ordering, and extensions to vision-language and diffusion model families. The signal remained robust under weight noise and sparse pruning, while an SVD-based weight-reduction variant was reported to be more robust than Frobenius cosine.