A Unified Framework for Evaluating the Trustworthiness of LLMs, AI Agents, and Multimodal Systems
Summary
Benchmark scores alone do not fully establish whether modern AI systems are trustworthy. This paper proposes a unified evaluation framework for large language models, agentic systems, and multimodal models while preserving assessment methods specific to each system. It connects output-level, trajectory-level, and cross-modal evaluation across eight dimensions: capability, robustness, safety, fairness, transparency, governance, oversight, and efficiency. Native metrics are mapped to common performance bands, with uncertainty estimates and traceable evidence intended to keep results interpretable. A meta-evaluation layer assesses whether the evaluation itself is valid, reliable, and reproducible. The framework also produces multidimensional profiles so strengths and weaknesses remain visible, and introduces safety-critical overrides to prevent aggregate scores from concealing serious failures. Links to governance frameworks, international standards, and European Union regulatory requirements are intended to connect technical assessment with oversight. The authors identify empirical validation across deployment contexts as an essential next step, so the framework is presented as a structured proposal rather than a completed validation of its effectiveness.