Researchers introduce DocHop, a benchmark for testing whether multimodal large language models can combine narrative context with evidence from multiple charts in document-style images. Unlike isolated chart or document question-answering evaluations, DocHop requires models to resolve a semantic reference from the text, identify relevant entities, and aggregate visual evidence through multi-step constraints. It contains 2,074 examples across six task categories. Human annotators exceed 90% accuracy, while the best evaluated model reaches 62.83%. Reasoning-enhanced models improve overall, but performance declines as compositional complexity increases.
AI News
The latest AI releases, research, products, and industry updates.
Loading...