Back to News
RSS feedarxiv.org

SIMLIFE Benchmarks Long-Horizon Pattern Understanding for Human-Agent Partnership

Summary

The paper introduces SimLife, a scalable platform for simulating long-term household life through rich visual observations, ground-truth action logs, and synthetic audio dialogues. Its benchmark, SimLife-BP, evaluates whether agents can infer latent behavioral rules from weeks or months of everyday observations rather than merely predict the next surface-level action. The benchmark includes 106 episodes averaging 15.49 hours and 38.57 in-game days, along with 1,439 question-answer pairs. Tasks cover direct, counterfactual, noisy, and inverse reasoning, with varying levels of rule hints. Evaluations of frontier models and architectures find that current systems often rely on frequency-based heuristics instead of constructing if-then rules from evidence. They also struggle when established behavioral patterns change. The authors identify long-context pattern understanding as a major bottleneck for future embodied agents and position SimLife as a platform for studying memory, personalization, adaptation, and long-horizon planning in everyday human-AI interaction.