RSS feedarxiv.org
SCAFFOLD Dataset Brings Structured Diagram Reasoning to Vision-Language Models
Summary
SCAFFOLD introduces structured training data for vision-language models that need to interpret diagrams in computer science papers. It contains figures paired with captions, context, questions, answers, and step-by-step reasoning traces. The largest version includes 157,387 pairs from 3,058 papers and 29,887 figures. Smaller 37K and 12K versions are also available, with baseline experiments conducted using Qwen2.5-VL-3B-Instruct.