Crab+ Unifies AV-LLMs, Reverses Negative Transfer

Crab+ introduces a novel approach to Audio-Visual Large Language Models, overcoming negative transfer via explicit cooperation in data and model design.

Crab+ Unifies AV-LLMs, Reverses Negative Transfer

The pursuit of unified scene understanding in multimodal intelligence hinges on effective Audio-Visual Large Language Models (AV-LLMs). However, conventional multi-task learning approaches for these models are plagued by substantial negative transfer, degrading performance on nearly 55% of tasks compared to single-task training. This stems from the inherent heterogeneity of audio-visual tasks, marked by varying granularity and distinct capability demands, leading to interference during joint training.

Tackling Task Heterogeneity with Explicit Cooperation

To overcome this critical bottleneck, the researchers introduce Crab$^{+}$, a scalable model designed for unified audio-visual scene understanding. Crab$^{+}$ tackles task heterogeneity through explicit cooperation at both the data and model levels. On the data front, they present AV-UIE v2, a comprehensive dataset featuring approximately 222K samples across 17 datasets and 7 tasks, augmented with explicit reasoning processes. This dataset enables the model to better capture cross-task relationships at diverse granularities. Complementing this, the model architecture features a unified interface to standardize task formulations and introduces Interaction-aware LoRA (I-LoRA). I-LoRA employs dynamic routing to explicitly model inter-task relationships, thereby coordinating distinct audio-visual interaction patterns and mitigating parameter interference.

Demonstrating Positive Transfer and Broad Applicability

The experimental results from Crab$^{+}$ are compelling. The model not only covers a broader range of tasks than existing unified AV-LLMs but also surpasses specialized models on various benchmarks. Crucially, Crab$^{+}$ reverses the negative transfer trend, achieving positive transfer where multi-task learning now outperforms single-task baselines in nearly 88% of evaluated tasks. These gains are consistent across diverse AV-LLM paradigms and are validated by in-depth visualizations, positioning Crab$^{+}$ as a robust advancement in holistic audio-visual scene understanding.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.