Visual TL;DR. Multi-modal memory bottleneck leads to Current benchmarks lacking. Current benchmarks lacking leads to M³Eval benchmark. M³Eval benchmark leads to Cognitive psychology grounded. M³Eval benchmark leads to Disentangled representation struggle. M³Eval benchmark leads to Interference patterns observed. M³Eval benchmark leads to Temporal grounding issues. M³Eval benchmark leads to Symbolic gaps identified.
- Multi-modal memory bottleneck: long-form video understanding requires robust information retention and recall
- Current benchmarks lacking: existing evaluations overlook systematic memory dimension assessment
- M³Eval benchmark: first comprehensive framework for probing multi-modal memory
- Cognitive psychology grounded: tasks designed to isolate key memory aspects
- Disentangled representation struggle: models struggle maintaining separate info from parallel video streams
- Interference patterns observed: models show weaknesses in separating interfering information
- Temporal grounding issues: challenges in recalling events in correct chronological order
- Symbolic gaps identified: difficulty connecting visual and textual information symbolically
Visual TL;DR