Back to News
RSS feedarxiv.org

Reward Exploits and Countermeasures in Conversational Humor

Summary

This study examines automated rewards for training language models to produce conversational humor, focusing on whether the rewards measure understandable surprise and predicted audience amusement or can be optimized through shortcuts. In controlled tests, an embedding-based surprise reward accepts word-shuffled replies almost as readily as genuinely witty responses. A fluency filter removes the shuffles, but the combined reward also rejects some witty replies and fails a later validation. A separate audience model is vulnerable to laughter cues placed in either speaker’s messages; normalizing those cues across speakers blocks the tested attacks, although unmatched expressions remain exploitable. The researchers conduct three reinforcement-learning runs with successive reward revisions. The final run raises the combined evaluation score by 0.0903 and cuts zero-score sessions by 40%, but its humor-specific improvement remains below the preregistered target. The findings show that reward countermeasures must suppress exploitable shortcuts without discarding the behavior the reward is intended to encourage.