Visual TL;DR. VLM computational demands leads to Existing pruning limitations. Existing pruning limitations solves GRIP-VLM framework. GRIP-VLM framework uses RL for discrete optimization. RL for discrete optimization employs GRPO paradigm. GRPO paradigm enables Direct discrete search. RL for discrete optimization enables Direct discrete search. Direct discrete search leads to Superior efficiency.
- VLM computational demands: escalating computational demands of VLMs driven by massive visual token processing
- Existing pruning limitations: existing training-aware pruning falters under aggressive compression due to approximations
- GRIP-VLM framework: novel framework for discrete vision-language model pruning
- RL for discrete optimization: formulates visual token pruning as a Markov Decision Process
- GRPO paradigm: Group Relative Policy Optimization augmented by supervised warm-up
- Direct discrete search: directly navigates the discrete search space for effective pruning decisions
- Superior efficiency: achieves unprecedented efficiency and adaptability in VLMs
Visual TL;DR
