MEA Uses Reward-Driven Multi-Agent Systems for Faithful Model Explanations
Summary
Machine-learning models used in high-stakes settings are often difficult for practitioners to interpret, while existing post-hoc explanation tools require expertise to configure and compare. The paper introduces MEA, a multi-agent framework in which a Proposer agent selects and configures explanation tools according to the question and modality, and an Actor agent converts their outputs into natural-language explanations optimized for faithfulness to the underlying model behavior. MEA covers tabular, text, and vision models and supports questions about feature attribution, counterfactual reasoning, and spurious features. Each question type is paired with a perturbation-based faithfulness metric. The authors report that frontier LLMs systematically generate unfaithful explanations. Across six datasets, MEA outperforms post-hoc explainers, agentic baselines, and closed-source baselines. Relative to the untrained backbone, reward-driven optimization improves faithfulness by 28% for tabular data, 21% for text, and 34% for vision, with an additional modality-adaptive penalty used during optimization. The findings suggest that agents can provide a scalable natural-language interface for ML explainability beyond fixed-purpose tools.