Back to News
RSS feedarxiv.org

The Attention Triangle Reveals Semantic Leakage in Audio-Video Diffusion Models

Summary

This paper examines how audio-video diffusion models coordinate text, sound, and visual content through three cross-attention pathways called the “attention triangle.” Its analysis finds that the audio-video pathway is bidirectional: audio can affect video generation, while video can affect audio generation. The pathway is shaped by biases encoded in the model parameters and can become a major source of semantic leakage. When a prompt conflicts with learned priors, these cross-modal interactions may override the intended conditioning and redirect semantics toward visually canonical but incorrect outcomes. The authors extract attention-based signals to show how semantic information is distributed and grounded across modalities. They also use the signals to induce leakage in controlled experiments, allowing them to isolate the contribution of individual routing interactions. Finally, the signals guide inference-time interventions that encourage more consistent alignment between modalities. Extensive experiments report improved semantic grounding while preserving generation quality.