Back to News
RSS feedarxiv.org

DriveHierarchy Benchmarks VLM Driving from Understanding to Execution

Summary

Evaluating VLM-based autonomous driving is difficult because driving competence combines grounding traffic participants and hazards, integrating information across views and time, anticipating how situations will evolve, and acting appropriately during interaction. The paper introduces DriveHierarchy, a hierarchical benchmark with four ranks: perceptual grounding, contextual memory, mental reasoning, and closed-loop execution. Its open-loop component unifies multiple open-source autonomous-driving datasets into 76,798 question-answer pairs covering 84,279 frames. The authors also build a closed-loop simulation platform with interactive scenario construction on a real-world road network and curate 100 scenarios for embodied evaluation. Experiments across 15 VLMs show structured, non-redundant differences in capability, a relationship between open-loop understanding and closed-loop driving, and a basis for diagnosing systems and guiding benchmark-based optimization. An anonymized project has been released on GitHub.