AdaCodec: Efficient Video MLLM Encoding

AdaCodec revolutionizes video MLLMs by using predictive visual coding to drastically cut tokenization costs and latency, achieving superior performance at a fraction of the budget.

Diagram illustrating AdaCodec's adaptive visual tokenization strategy for video MLLMs.
AdaCodec's adaptive approach to visual tokenization.
Visual TL;DR
Video MLLM InefficiencyDriver
processing adjacent frames as independent images leads to redundant tokens
From the articleThe inherent temporal redundancy in video, where adjacent frames largely overlap, presents a fundamental inefficiency for current video multimodal large language models (video MLLMs).
Temporal RedundancyContext
adjacent video frames largely overlap, causing inflated computational costs
From the articleThe inherent temporal redundancy in video, where adjacent frames largely overlap, presents a fundamental inefficiency for current video multimodal large language models (video MLLMs).
AdaCodec IntroducedCore
a new dynamic and efficient video interface for MLLMs
From the article 3 mentionsInstead of encoding every frame fully, this system, instantiated as AdaCodec, selectively transmits a full reference frame only when scene prediction is unreliable.
Predictive Visual CodingCore
intelligently manages visual token transmission based on scene prediction
From the articleThe core innovation lies in a 'predictive visual code' that intelligently manages visual token transmission.
Selective Frame EncodingContext
From the articleInstead of encoding every frame fully, this system, instantiated as AdaCodec, selectively transmits a full reference frame only when scene prediction is unreliable.
Compact P-tokensCore
From the articleOtherwise, it encodes inter-frame changes, encompassing motion and prediction residuals, using compact 'P-tokens'.
Reduced Token CountEffect
significantly minimizes visual tokens required for video understanding
From the articleEven at a drastically reduced token budget (1/7th), AdaCodec with 32k tokens outperforms the 224k baseline on all long-video benchmarks.
Efficiency GainsOutcome
drastically cuts tokenization costs and latency for video MLLMs
From the articleThis efficiency leap makes real-time video analysis and interaction far more feasible.
Superior PerformanceOutcome
achieves better results at a fraction of the computational budget

The inherent temporal redundancy in video, where adjacent frames largely overlap, presents a fundamental inefficiency for current video multimodal large language models (video MLLMs). These models typically process each sampled frame as an independent image, leading to redundant visual tokens and inflated computational costs. A new approach, detailed on arXiv, challenges this paradigm by proposing a more dynamic and efficient video interface.

Predictive Visual Coding for Reduced Redundancy

The core innovation lies in a 'predictive visual code' that intelligently manages visual token transmission. Instead of encoding every frame fully, this system, instantiated as AdaCodec, selectively transmits a full reference frame only when scene prediction is unreliable. Otherwise, it encodes inter-frame changes, encompassing motion and prediction residuals, using compact 'P-tokens'. This adaptive strategy significantly minimizes the number of visual tokens required for video understanding.

Substantial Gains in Efficiency and Performance

AdaCodec demonstrates marked improvements over the baseline Qwen3-VL-8B model across eleven benchmarks. Even at a drastically reduced token budget (1/7th), AdaCodec with 32k tokens outperforms the 224k baseline on all long-video benchmarks. Furthermore, for general-video benchmarks, it not only elevates average scores but also slashes the time-to-first-token from 9.26s to a mere 1.62s. This efficiency leap makes real-time video analysis and interaction far more feasible.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Daniel Singer

Written by

Daniel Singer

Editor, StartupHub.ai

Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.

More from Daniel Singer