RLCE Uses Comparative Evidence to Train Socially Intelligent AI
Summary
Socially intelligent AI is difficult to train without extensive human annotations because affect, intent, preference, and pragmatic meaning are often ambiguous, and unlike mathematics or coding, social predictions lack reliable verification oracles. The paper introduces Reinforcement Learning with Comparative Evidence (RLCE), which learns from unlabeled training data without constructing rewards from ground-truth annotations. For distinct answers produced in a rollout group, RLCE creates tests for observable evidence favoring one answer over another, validates those tests against the input, and aggregates the results to identify the best-supported interpretation. The tests are regenerated as the policy produces new answers, allowing the evidence criteria to evolve with the policy. Across four benchmarks covering affect, pragmatics, communicative intent, and preference, RLCE achieves the strongest performance among seven methods that use no ground-truth training labels for rewards, including consensus, policy LLM-judge verification, multimodal co-evolution, and rubric-based rewards. Its gains over the strongest baseline reach 18.93 points. Further analyses report greater separation in reward variation between correct and incorrect predictions than with the compared rubric methods, show that RLCE can reverse erroneous policy-derived preferences, and find benefits from pairwise test construction, compositional aggregation, and on-policy test evolution.