MMDiff: Auditing and Steering MLLMs

MMDiff, a new multimodal model-diffing framework, enables granular control and understanding of MLLMs by isolating and manipulating specific behavioral features.

Diagram illustrating the MMDiff framework for multimodal model-diffing and feature control in MLLMs.
MMDiff provides a structured approach to understanding and manipulating the internal workings of multimodal large language models.
Visual TL;DR
MLLMs are OpaqueDriver
From the article 5 mentionsThe opaque nature of Multimodal Large Language Models (MLLMs) presents a significant hurdle for researchers and developers seeking to understand, audit, or control their complex behaviors.
SAEs Fall ShortDriver
sparse autoencoders lack ability to pinpoint features altered by multimodal training
From the articleWhile techniques like Sparse Autoencoders (SAEs) offer post-hoc analysis, they often fall short in pinpointing features altered by multimodal training or enabling targeted interventions.
MMDiff FrameworkCore
From the article 9+ mentionsAddressing this gap, a new framework called MMDiff multimodal model-diffing emerges as a critical tool for dissecting and manipulating MLLM capabilities.
Trains Multimodal SAEsContext
identifies specific features modified during the multimodal adaptation process
From the article 3 mentionsMMDiff introduces a novel approach by training multimodal SAEs.
Isolates BehaviorsEffect
enables granular control and understanding by isolating specific behavioral features
From the article 4 mentionsThe results are compelling: MMDiff successfully isolates sparse, causally specific features.
Causal ControlEffect
transforms identified features into actionable interfaces for targeted interventions
From the article 5 mentionsSecondly, MMDiff facilitates task-specific feature detection through per-token contrastive firing analysis, pinpointing the causal features responsible for particular behaviors.
Steer MLLMsOutcome
provides a critical tool for auditing and steering multimodal large language models
From the article 6 mentionsThe ability to steer features offers a pathway to enhance model capabilities, pushing accuracy metrics higher, while simultaneously improving safety by mitigating vulnerabilities.
Contents(3)

The opaque nature of Multimodal Large Language Models (MLLMs) presents a significant hurdle for researchers and developers seeking to understand, audit, or control their complex behaviors. While techniques like Sparse Autoencoders (SAEs) offer post-hoc analysis, they often fall short in pinpointing features altered by multimodal training or enabling targeted interventions. Addressing this gap, a new framework called MMDiff multimodal model-diffing emerges as a critical tool for dissecting and manipulating MLLM capabilities.

Unpacking Multimodal Behaviors with MMDiff

MMDiff introduces a novel approach by training multimodal SAEs. This allows for the identification of specific features that are modified during the multimodal adaptation process when a base language model is fused with visual understanding. The framework's core innovation lies in its ability to transform these identified features into actionable interfaces. This enables researchers to not only discover what changes multimodal training induces but also to directly control these specific behavioral aspects.

From Isolation to Causal Control

The utility of MMDiff is demonstrated across three key applications. Firstly, feature isolation is achieved by comparing SAEs from a base language model against those from its multimodal counterpart. This diffing process highlights the precise alterations introduced by multimodal data. Secondly, MMDiff facilitates task-specific feature detection through per-token contrastive firing analysis, pinpointing the causal features responsible for particular behaviors. Finally, and perhaps most impactfully, the framework enables feature-level control. This allows for the causal removal or steering of discovered feature directions, offering a direct mechanism to modify model outputs.

The researchers validated MMDiff across three prominent MLLM families LLaVA-MORE, PaliGemma 2, and InternVL3.5. Evaluations focused on visual-spatial understanding, multimodal safety, and Optical Character Recognition (OCR). The results are compelling: MMDiff successfully isolates sparse, causally specific features. Removing these features selectively degraded spatial understanding by an average of 12% and OCR accuracy by 17%. Crucially, on multimodal safety benchmarks, MMDiff reduced attack success rates by 24%, all without impacting general Visual Question Answering (VQA) performance. Furthermore, steering these identified features improved spatial and OCR accuracy by +3.6% and +1.8% respectively, outperforming a standard single-layer steering baseline.

Strategic Implications for MLLM Development

These findings position MMDiff multimodal model-diffing as more than just an interpretability tool. It serves as a powerful mechanism for auditing MLLM behavior, a critical step for deploying these models in sensitive applications. The ability to steer features offers a pathway to enhance model capabilities, pushing accuracy metrics higher, while simultaneously improving safety by mitigating vulnerabilities. This research provides a tangible method for achieving more trustworthy and performant MLLMs, addressing a core challenge in the field.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.