CapMem studies whether textual captions can serve as reusable episodic memory for wearable assistants that need to reason over long egocentric videos. The work addresses practical limits of vision-language models, including bounded frame budgets, rising visual-token costs, and retrieval failures in long contexts. It introduces the Episodic Memory Video Caption QA task and a human-annotated benchmark containing 75 videos totaling 33.7 hours, with 1,000 multiple-choice questions covering 16 scenarios. For videos longer than 20 minutes, full-coverage CaptionQA using 30-second caption windows outperformed direct VideoQA on 10 of 12 models; 60-second windows did so on 8 of 12 models. A matched-frame control across six Qwen models still showed mean accuracy gains of 3.22 and 2.55 points for the two window sizes, respectively, indicating that the result was not explained solely by using more frames. The authors’ caption-guided retrieve-and-verify system improved accuracy by up to 5.3 points. The findings support caption memory as a practical approach for episodic reasoning over long first-person video, while the paper’s evidence is specifically based on the introduced benchmark and reported model evaluations.
AI News
The latest AI releases, research, products, and industry updates.
Loading...