Back to News
RSS feedarxiv.org

RESCUE-BENCH Tests LLMs on Relationship-Aware Emotional Support

Summary

The paper introduces relation-aware emotional support conversation as a task for testing whether large language models can track changing interpersonal dynamics and use them to provide more effective support. It presents RESCUE, a benchmark built from real couple and family interview conversations, with 191 samples, 7,079 annotated turns, and 1,064.8 minutes of video. Its annotations cover socio-emotional and support-related dynamics, and the benchmark defines six tasks across relational understanding and relation-sensitive support. Experiments with ten LLMs show relatively strong performance on tasks based on local emotional or intervention cues. However, the models struggle with relation-intensive tasks, including relation pattern prediction, viewpoint prediction, and support strategy prediction. The findings point to limitations in current LLMs’ ability to model interpersonal relationships and make support decisions that account for those relationships.