GeoOutageBench Introduces a Benchmark for Ambiguity-Aware Multimodal Power-Outage KGQA
Summary
GeoOutageBench is a benchmark for evaluating large-language-model-based geospatiotemporal question answering over multimodal knowledge graphs for power-outage and resilience analysis. Its knowledge graph combines visual, textual, and structured information from outage records, remote-sensing data, weather observations, storm and power events, geographic entities, and domain ontologies. The benchmark organizes competency queries by difficulty, covering spatiotemporal containment and proximity, spatiotemporal co-occurrence, multimodal evidence, and hypothetical evaluation. It supports configurable assessment of three related capabilities: interpreting ambiguous natural-language questions as SPARQL queries, measuring the usefulness of ontologies, and evaluating the accuracy of multimodal knowledge-graph retrieval and question answering. The authors present it as a foundation for assessing LLM-KG systems intended to support real-world infrastructure resilience analysis. The benchmark, source code, data, results, and documentation are available through the project’s GitHub repository.