Multi-Objective Explanation Tuning (METRIX)
• Created a novel method that jointly optimizes infidelity and sensitivity in model interpretability (ImageNet, MNIST). • Achieved 99.9% lower sensitivity than baseline without sacrificing model faithfulness; visualized 200+ class explanations. • First to propose a closed-form update rule for infidelity-sensitivity balancing in feature attribution.
Open project