Visual TL;DR. MLLMs are Opaque leads to SAEs Fall Short. MLLMs are Opaque addresses MMDiff Framework. SAEs Fall Short gap filled by MMDiff Framework. MMDiff Framework by Trains Multimodal SAEs. Trains Multimodal SAEs allows Isolates Behaviors. Isolates Behaviors enables Causal Control. Causal Control results in Steer MLLMs. Isolates Behaviors to Steer MLLMs.
- MLLMs are Opaque: complex behaviors of multimodal large language models are hard to understand or control
- SAEs Fall Short: sparse autoencoders lack ability to pinpoint features altered by multimodal training
- MMDiff Framework: new multimodal model-diffing framework for dissecting and manipulating MLLM capabilities
- Trains Multimodal SAEs: identifies specific features modified during the multimodal adaptation process
- Isolates Behaviors: enables granular control and understanding by isolating specific behavioral features
- Causal Control: transforms identified features into actionable interfaces for targeted interventions
- Steer MLLMs: provides a critical tool for auditing and steering multimodal large language models
Visual TL;DR
