ROAR Unifies Runs Across Heterogeneous AI Research Systems
Summary
AI-driven research systems often produce expensive runs that are difficult to compare because teams and frameworks store results in incompatible formats. The paper introduces ROAR, an infrastructure designed to reconcile these outputs and support analysis across runs with different objectives and scoring functions. Its relational schema and parsing layer preserve data lineage and temporal structure while normalizing results, and can accommodate new systems without changing the schema. The authors build a corpus of more than 900 runs from multiple AI-driven research systems. Analysis of the pooled corpus confirms that identical configurations can reach different scores, shows that many runs obtain most of their gains early, and finds that the value of incorporating prior solutions varies by problem. ROAR is also used to configure new runs, indicating that the corpus can guide system operation rather than serve only as a retrospective database. The authors present it as infrastructure for uncovering problem-dependent search behavior that remains difficult to see when runs are siloed.