Keep It CALM: Limits of Global Safety Signals in Text-to-Image Generation
Summary
The paper examines training-free safeguards for text-to-image generation that apply a reusable global safety signal across prompts. Its geometric analysis identifies a consistent coverage-selectivity trade-off: compact unsafe subspaces do not cover heterogeneous unsafe semantics, while broader aggregation increasingly distorts benign prompts that are close to unsafe concepts. The authors propose CALM, or Counterfactual Adaptive Local Modulation, as a training-free alternative based on prompt-local correction. Using matched unsafe-benign anchors, CALM identifies the unsafe categories active in each prompt and minimally edits only violating token representations toward the safe side. It also suppresses unsafe residual components that are positively aligned with the prompt representation. In broad evaluation, the method improves suppression of unsafe content while preserving benign utility. The results support local counterfactual correction as a more selective approach than uniform removal of a global unsafe signal.