Back to News
RSS feedarxiv.org

Build2SPARQL Introduces a Large-Scale Text-to-SPARQL Benchmark for Building Knowledge Graphs

Summary

Build2SPARQL is a large-scale benchmark for translating natural-language questions into SPARQL queries over building automation knowledge graphs. The dataset targets graphs built with the Brick and ASHRAE 223P ontologies, which provide a machine-readable basis for AI applications and language-agent interfaces. Its knowledge-graph-grounded pipeline generates and validates SPARQL entirely through graph-traversal code, while large language models are used only to phrase the natural-language questions, separating query correctness from model behavior. The benchmark covers six query-pattern families: linear chains, branching, UNION, aggregation, OPTIONAL, and attribute-filtered queries, with each query expressed in five vocabulary registers. Applied to 201 building knowledge graphs, including 180 Brick graphs and 21 ASHRAE 223P graphs, the pipeline produced 6,136 executable queries and 30,680 questions. In a two-rater review of 300 questions, 98.8% were judged semantically faithful, 98.8% natural, and 84.0% operationally plausible. A retrieval-augmented evaluation using three open-weight language models improved exact-match accuracy from 0.2-20% in zero-shot settings to 56-65% with three retrieved examples, indicating that the benchmark can support systematic evaluation of building-graph querying systems.