Back to News
RSS feedarxiv.org

SCAFFOLD Dataset Brings Structured Diagram Reasoning to Vision-Language Models

Summary

SCAFFOLD introduces structured training data for vision-language models that need to interpret diagrams in computer science papers. It contains figures paired with captions, context, questions, answers, and step-by-step reasoning traces. The largest version includes 157,387 pairs from 3,058 papers and 29,887 figures. Smaller 37K and 12K versions are also available, with baseline experiments conducted using Qwen2.5-VL-3B-Instruct.