Back to News
RSS feedarxiv.org

Knowledge Graph Framework Evaluates Whether LLMs Understand Context

Summary

The paper asks whether large language models genuinely understand context or mainly rely on large-scale pattern matching. It argues that common measures such as BLEU and perplexity capture surface-level performance but do not adequately assess whether answers are grounded in the supplied context. To address this gap, the authors introduce a knowledge graph-based evaluation framework for contextual understanding in question answering. Its central metric, Semantic Structural Similarity for KGs (S3KG), combines structural and semantic signals into one score. The framework also includes a diagnostic analysis that identifies and categorizes reasoning errors at the knowledge-graph triplet level, allowing model failures to be examined more precisely. Across nine benchmarks, S3KG reports F1 improvements of up to 7.6 points over the strongest baseline and reaches an AUROC of up to 0.973. The results position the framework as a way to evaluate context grounding and diagnose reasoning failures beyond conventional surface-oriented metrics.