Videos

Bin Yu - Not All Interpretation Methods Are Created Equal: Principled Evaluation of AI Interpretability

September 1, 2026
Abstract
As AI transitions to agentic sociotechnical deployments, uninterpretable models pose serious threats to reliability, public trust, and contestability. However, not all interpretation methods are created equal: for interpretability to enhance deployment safety and effectiveness, interpretation methods must be evaluated as a whole, rather than validated through isolated interpretation instances. In this paper, we advocate elevating the evaluation of interpretation methods as a core assessment metric for AI system deployment, with clearly stated principles. To this end, we propose TRUST, a principled framework that builds upon Veridical Data Science (VDS) and the Predictive, Descriptive and Relevance (PDR) foundations. To differentiate interpretation methods effectively, TRUST systematically grades interpretation methods across five progressive tiers: Task-Aligned PDR Foundations, Rigorous Epistemic Grounding, Unifying Semantic Understanding, Scope Transferability, and Targeted Prospective Validation. We present a start-to-finish case study of the iterative Random Forest (iRF) family under TRUST and contrast with other interpretation methods that fail to clear specific tiers. Moreover, we will briefly make connections among contestability, TRUST, and our Green Shielding framework (an user-centric approach to trustworthy AI). Ultimately, TRUST principles foster contestability and epistemic grounding across both human-machine and fully autonomous agentic workflows.