#Interpretability
4 articles with this tag

AI Research
MMDiff: Auditing and Steering MLLMs
MMDiff, a new multimodal model-diffing framework, enables granular control and understanding of MLLMs by isolating and manipulating specific behavioral features.
about 3 hours ago

AI Research
Hybrid AI Models Get Orthogonal
OrthoReg, a novel regularization method, ensures clear separation between symbolic and neural components in hybrid dynamical systems, boosting interpretability and generalization.
about 2 months ago
AI Research
Phase Dominance in AI Image Recognition
AI image classifiers exhibit a striking phase dominance for identity encoding, mirroring human vision principles, with architectural differences shaping its expression.
about 2 months ago

AI Research
Certified Circuits for Stable AI Explanations
New 'Certified Circuits' framework provides provable stability for AI model explanations, yielding more accurate and compact circuits.
5 months ago