TwelveLabs Builds Video Memory Layer
James Le of TwelveLabs discusses the critical need for memory in video AI, detailing the limitations of current systems and introducing their 'Jockey' platform as a solution.

Visual TL;DR
systems treat video as 'bag of frames' or transcript, losing spatiotemporal context
From the article 8 mentionsJames Le, speaking at the AI Engineering World's Fair, introduced the concept of a 'memory layer' for video intelligence, arguing that current AI systems fail to truly understand video due to a lack of memory and a misapplication of language model principles.
forcing video into text sequences discards crucial spatiotemporal relationships and event definitions
From the article 2 mentionsWrong Context: Forcing video into a text-based sequence by sampling frames or extracting transcripts loses the spatiotemporal relationships that define events.
introduces a 'memory layer' for video intelligence to overcome current AI limitations
From the article 3 mentionsJockey is currently in private beta, and developers working on content assembly or organization workflows are encouraged to explore the platform.
enables true understanding of video by retaining continuity and relationships within spatiotemporal volume
From the article 9 mentionsHe outlined five key principles for building a video memory layer:
Jockey uses these principles to build a robust and comprehensive video memory system
From the articleLe proposed representing video collections as a 'context graph,' a durable, queryable representation connecting moments, entities, appearances, relationships, and corpus-level context.
AI can now grasp complex events and relationships, not just isolated frames or words
From the article 3 mentionsThe continuity and relationships within this volume are crucial for understanding, yet current approaches often treat video as a stack of images or a transcript, discarding this vital context.
paving the way for more sophisticated and human-like interaction with video content
From the articleJames Le, speaking at the AI Engineering World's Fair, introduced the concept of a 'memory layer' for video intelligence, arguing that current AI systems fail to truly understand video due to a lack of memory and a misapplication of language model principles.
Contents(6)
© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.
Written by
Daniel SingerEditor, StartupHub.ai
Daniel Singer is the editor of StartupHub.ai, a technology expert and thought leader on AI and its applications across sectors, from fintech and healthcare to developer tooling and consumer software. He writes and tests the tools covered here thoroughly and regularly, and built StartupHub.ai to give founders, operators and buyers a clearer read on what they are actually being sold.
More from Daniel Singer