AdaCodec: Efficient Video MLLM Encoding
AdaCodec revolutionizes video MLLMs by using predictive visual coding to drastically cut tokenization costs and latency, achieving superior performance at a fraction of the budget.
Visual TL;DR
processing adjacent frames as independent images leads to redundant tokens
From the articleThe inherent temporal redundancy in video, where adjacent frames largely overlap, presents a fundamental inefficiency for current video multimodal large language models (video MLLMs).
adjacent video frames largely overlap, causing inflated computational costs
From the articleThe inherent temporal redundancy in video, where adjacent frames largely overlap, presents a fundamental inefficiency for current video multimodal large language models (video MLLMs).
a new dynamic and efficient video interface for MLLMs
From the article 3 mentionsInstead of encoding every frame fully, this system, instantiated as AdaCodec, selectively transmits a full reference frame only when scene prediction is unreliable.
intelligently manages visual token transmission based on scene prediction
From the articleThe core innovation lies in a 'predictive visual code' that intelligently manages visual token transmission.
From the articleInstead of encoding every frame fully, this system, instantiated as AdaCodec, selectively transmits a full reference frame only when scene prediction is unreliable.
From the articleOtherwise, it encodes inter-frame changes, encompassing motion and prediction residuals, using compact 'P-tokens'.
significantly minimizes visual tokens required for video understanding
From the articleEven at a drastically reduced token budget (1/7th), AdaCodec with 32k tokens outperforms the 224k baseline on all long-video benchmarks.
drastically cuts tokenization costs and latency for video MLLMs
From the articleThis efficiency leap makes real-time video analysis and interaction far more feasible.
achieves better results at a fraction of the computational budget
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer