Robust Signaling Under an Informed Adversary Collapses onto Salience
Summary
The paper studies a forced-choice signaling game in which an informed adversary knows the target, observes the signal, and uses a persuasion budget, beta, to argue for the strongest wrong answer. It asks how the signal that best protects the truth changes when the adversary can influence the shared audience. On 108 confirmatory items, the adversary-robust optimum exactly matches the salience pole identified in earlier work. In a 200,000-item pool, the two choices differ on only 2,748 items, and those disagreements occur exactly where the prior salience-to-Bayes coordinate is undefined. Where that coordinate is defined, increasing the adversary's budget moves the optimum entirely from Bayesian discrimination toward salience; at beta = 0, the game reproduces the earlier oracle model with listener temperature tau = 1. The shift is not confined to a boundary case: 18.2% of the pool has an optimum that changes under a finite budget, with an exact critical budget for each item. The authors also test two adversary framings with seven language models on 108 items. The framings change the selected option on 30 to 77 items, despite an exact no-effect rate used for comparison. However, because the adversary-aware target and the salience target are identical on the relevant items, the measurements cannot establish whether models moved toward adversary awareness or simply toward salience. The authors describe this as a structural identification limit rather than evidence of no effect. They recommend checking whether a robust target coincides with a heuristic target before using an evaluation to measure adversary awareness.